Back to feed
arXiv cs.CL·

The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer

Signal
78
Hype
15
In three linesPruned models pass multiple-choice benchmarks but fail in open generation. Multilingual study shows that under high-sparsity pruning (Wanda), correct answers are demoted rather than erased: they reappear with beam search or sampling. Multiple-choice benchmarks overstate the usability of compressed LLMs.
Read source
Your take?
BenchmarksEvalsFine-tuning

Summary generated by Claude — human-verified