Back to feed
arXiv cs.CL·

Replay What Matters: Off-Policy Replay for Efficient LLM Reinforcement Unlearning

Signal
78
Hype
15
In three linesReRULE improves LLM unlearning via off-policy replay for hard cases. The method stores low-reward rollouts near the forget/retain boundary in a replay buffer and reuses them through importance-sampled updates. On MUSE-Books, it increases Retain Quality from 46.3 to 56.2 with +5–11% training overhead.
Read source
Your take?
Reinforcement learningAI safetyAlignmentPapers

Summary generated by Claude — human-verified