SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Signal
78
Hype
25
In three linesSparse Autoencoders (SAEs) decompose activations into interpretable features, but this study shows that clamping a 'harmful' feature does not eliminate the behavior—it can recover via other residual pathways. Even with active intervention, 95.8% behavior recovery is achievable in refusal-steering, exposing a gap between feature-level control and behavioral completeness.Read source
Your take?
Summary generated by Claude — human-verified