The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge
Signal
72
Hype
15
In three linesStudy of relationship between internal confidence (log-probabilities), external evaluation (LLM-as-judge), and final accuracy in multi-agent debate. Framework with Constructor and Auditor reveals role asymmetry: Constructor's confidence predicts judged reasoning quality 2× better and detects critical failures more reliably (AUROC 0.804 vs 0.634 for Auditor).Read source
Your take?
Summary generated by Claude — human-verified