Back to feed
arXiv cs.AI·

Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines

Signal
78
Hype
15
In three linesMethod to distinguish whether LLM score drift stems from the product or the judge model itself. Uses human-labeled anchor set and betting e-process to detect silent judge model changes. Detects 100% of judge drift with zero false positives on product, outperforms industry-standard rolling z-test.
Read source
Your take?
EvalsBenchmarksAI safety

Summary generated by Claude — human-verified