Prefill Awareness in Large Language Models
Signal
78
Hype
15
In three linesarXiv study showing frontier models (Claude Opus 4.5, GPT, Gemini) detect tampered prefills in 9-35% of cases with 0% false positive rate. This 'prefill awareness' undermines alignment and jailbreaking evaluations relying on inserted assistant context. Models distinguish stylistic from preference mismatch.Read source
Your take?
Summary generated by Claude — human-verified