Back to feed
arXiv cs.CL·

CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning

Signal
78
Hype
15
In three linesCoRA aligns model confidence with chain-of-thought rationale quality. A GRPO-based RL framework jointly rewards answer correctness, committed-answer probability, and rubric-based rationale support. On MedQA, MathQA, OpenBookQA: 26.51% reduction in confidence-rationale alignment error across three open-weight LLMs.
Read source
Your take?
ReasoningReinforcement learningEvalsAlignment

Summary generated by Claude — human-verified