Back to feed
arXiv cs.AI·

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

Signal
78
Hype
15
In three linesRealMath-Eval, a benchmark of 224 real exam responses, reveals that state-of-the-art LLM judges fail to evaluate authentic human reasoning (MSE ~2.96 vs ~1.17 on synthetic solutions). Analysis shows human errors form a more diverse error space than synthetic errors, with higher information-theoretic surprisal.
Read source
Your take?
EvalsBenchmarksReasoning

Summary generated by Claude — human-verified