Back to feed
arXiv cs.CL·

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

Signal
72
Hype
15
In three linesStudy of relationship between internal confidence (log-probabilities), external evaluation (LLM-as-judge), and final accuracy in multi-agent debate. Framework with Constructor and Auditor reveals role asymmetry: Constructor's confidence predicts judged reasoning quality 2× better and detects critical failures more reliably (AUROC 0.804 vs 0.634 for Auditor).
Read source
Your take?
Multi-agentAI AgentsReasoningEvalsPapers

Summary generated by Claude — human-verified