Back to feed
arXiv cs.AI·

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

Signal
78
Hype
25
In three linesLLM judges used to evaluate AI models are unstable under post-decision interaction. On MT-Bench and AlpacaEval, researchers show initial judgments can be reversed through targeted challenges, degrading agreement with human preferences and shifting benchmark rankings. They introduce the Evaluation Robustness Score (ERS) to measure this fragility.
Read source
Your take?
EvalsBenchmarksAI safetyAlignment

Summary generated by Claude — human-verified