Back to feed
arXiv cs.LG·

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

Signal
72
Hype
15
In three linesStudy on evaluating reasoning models beyond accuracy alone. Authors introduce two metrics: susceptibility (whether bias breaks a previously correct answer) and acknowledgment (whether the trace explicitly references injected biased content). On GSM8K, GPT-4o and Claude Sonnet 4 show similar susceptibility rates (1.3% vs 1.2%) but substantially different acknowledgment rates (13.0% vs 75.0%).
Read source
Your take?
EvalsReasoningAI safetyAlignment

Summary generated by Claude — human-verified