Back to feed
arXiv cs.CL·

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

Signal
72
Hype
15
In three linesStudy evaluating 42 LLMs (proprietary and open-source) on their ability to measure item discrimination in reading comprehension. Models fail: Spearman correlation of 0.152 in direct prediction, 0.241 in CTT calibration. LLMs do not reliably capture how assessment items distinguish students of different proficiency levels.
Read source
Your take?
BenchmarksEvalsPapers

Summary generated by Claude — human-verified