Back to feed
arXiv cs.LG·

SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector

Signal
75
Hype
15
In three linesSAGE is a post-hoc method to improve selective unlearning in LLMs. It corrects final update vectors by suppressing components damaging retention, without rerunning the original unlearning pipeline. Tested across multiple methods and scales, SAGE reduces the forget-retain trade-off.
Read source
Your take?
AlignmentPapers

Summary generated by Claude — human-verified