Back to feed
arXiv cs.AI·

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

Signal
78
Hype
15
In three linesResearchers demonstrate that aligned LLMs remain vulnerable to token injections at any generation step, not just early tokens. Alignment with internal refusal directions does not predict robustness. Training directly on perturbed generation trajectories improves resistance to mid-sequence attacks.
Read source
Your take?
AI safetyAlignmentReasoning

Summary generated by Claude — human-verified