Back to feed
arXiv cs.AI·

Prefill Awareness in Large Language Models

Signal
78
Hype
15
In three linesarXiv study showing frontier models (Claude Opus 4.5, GPT, Gemini) detect tampered prefills in 9-35% of cases with 0% false positive rate. This 'prefill awareness' undermines alignment and jailbreaking evaluations relying on inserted assistant context. Models distinguish stylistic from preference mismatch.
Read source
Your take?
AI safetyAlignmentEvalsClaude

Summary generated by Claude — human-verified