Back to feed
arXiv cs.AI·

Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking

Signal
72
Hype
15
In three linesNew method to detect implicit reward hacking in LLMs without a reward model. The "self-commitment latency" metric measures how early a reasoning context commits to the model's final answer. Tested on Qwen2.5-3B with GSM8K: AUROC 0.878 for detecting prompt-shortcut contexts.
Read source
Your take?
ReasoningAI safetyAlignmentEvals

Summary generated by Claude — human-verified