The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models
Signal
78
Hype
15
In three linesStereological theory of LLM benchmark coverage. For d_eff ∈ [2.86, 4.80], structural blind spot exceeds runner-up score gap by two orders of magnitude. Submodular greedy algorithm identifies 4 stable benchmarks; 7 of 12 suffice for 90% coverage. Validation across 12 internal benchmarks and 27 Chatbot Arena categories.Read source
Your take?
Summary generated by Claude — human-verified