Back to feed
arXiv cs.CL·

Output Vector Editing for Memorization Mitigation in Large Language Models

Signal
78
Hype
15
In three linesMemorization suppression method in LLMs via output vector editing of MLP neurons. Tested on 4 models (360M-7B parameters), achieves 87.9% suppression on OLMo-7B with 6831 memorized sequences. Complementary approach to existing neuron ablation methods.
Read source
Your take?
AI safetyAlignmentPapersBenchmarks

Summary generated by Claude — human-verified