Back to feed
arXiv cs.AI·

Evaluation of LLMs for Mathematical Formalization in Lean

Signal
78
Hype
15
In three linesComparison of LLMs for generating formal proofs in Lean 4. Gemini 3.1 Pro and Claude Opus 4.7 achieve best performance (92% and 86% success rates respectively via refine@32). NVIDIA Nemotron 3 Super and GPT-OSS 120B offer best cost-efficiency (<$0.01 per correct proof).
Read source
Your take?
BenchmarksClaudeGeminiReasoningPapers

Summary generated by Claude — human-verified