CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning
Signal
78
Hype
15
In three linesCoRA aligns model confidence with chain-of-thought rationale quality. A GRPO-based RL framework jointly rewards answer correctness, committed-answer probability, and rubric-based rationale support. On MedQA, MathQA, OpenBookQA: 26.51% reduction in confidence-rationale alignment error across three open-weight LLMs.Read source
Your take?
Summary generated by Claude — human-verified