Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking
Signal
72
Hype
15
In three linesNew method to detect implicit reward hacking in LLMs without a reward model. The "self-commitment latency" metric measures how early a reasoning context commits to the model's final answer. Tested on Qwen2.5-3B with GSM8K: AUROC 0.878 for detecting prompt-shortcut contexts.Read source
Your take?
Summary generated by Claude — human-verified