Back to feed
arXiv cs.CL·

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

Signal
72
Hype
25
In three linesLLMs become overconfident when reasoning budget exceeds a task-specific threshold. This phenomenon, termed Calibration Drift Under Reasoning (CDUR), is studied on Llama-3.1-8B and Llama-3.3-70B. Authors propose CABStop, a calibration-aware stopping rule that halts reasoning when confidence diverges from actual accuracy.
Read source
Your take?
LlamaReasoningEvalsAI safetyAlignment

Summary generated by Claude — human-verified