Back to feed
arXiv cs.AI·

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

Signal
72
Hype
15
In three linesComparison of two intervention methods on safety refusals in chat models: Difference-in-Means (DiM) vs Iterative Nullspace Projection (INLP). On 5 open-weight models, INLP counterfactual flipping matches DiM for refusal suppression, while nullspace projection is consistently weaker. The two approaches operate differently in activation space.
Read source
Your take?
AI safetyAlignmentPapersOpen source

Summary generated by Claude — human-verified