Back to feed
arXiv cs.AI·

MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis

Signal
78
Hype
15
In three linesMA-ProofBench is the first formal theorem-proving benchmark dedicated to Mathematical Analysis with 200 formalized theorems across two difficulty levels (undergraduate and Ph.D.). GPT-5.5 achieves only 16% Pass@8 on Level I and 5% on Level II, exposing major gaps in LLMs' advanced formal reasoning capabilities.
Read source
Your take?
BenchmarksReasoningGPT

Summary generated by Claude — human-verified