Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval
Signal
75
Hype
15
In three linesParaEval, an evaluation framework, corrects MCQA benchmark biases by testing models with multiple paraphrases per answer option. Standard scores conflate familiarity with exact phrasing and actual understanding. On 1B-8B models with identical knowledge, ParaEval reduces artificial performance gaps from 2 points to below 1 point.Read source
Your take?
Summary generated by Claude — human-verified