Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge
Signal
78
Hype
15
In three linesJudge-LS evaluates whether LLMs used as automatic judges exhibit language bias. On 419 LLMBar benchmark items transformed into English, Chinese, and mixed-language variants, models show 10.7–14.4% preference flips across languages, with highest accuracy in English. Translation-equivalent probes reveal no systematic English preference, though most are judged as ties.Read source
Your take?
Summary generated by Claude — human-verified