Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories
Signal
78
Hype
15
In three linesResearchers demonstrate that aligned LLMs remain vulnerable to token injections at any generation step, not just early tokens. Alignment with internal refusal directions does not predict robustness. Training directly on perturbed generation trajectories improves resistance to mid-sequence attacks.Read source
Your take?
Summary generated by Claude — human-verified