Back to feed
arXiv cs.LG·

Natively Unlearnable Large Language Models

Signal
78
Hype
25
In three linesNULLs (Natively Unlearnable LLMs) is an architecture that isolates each data source's contributions in distinct parameters (sinks) while maintaining a shared backbone. Tested on ~6M Wikipedia articles, it enables unlearning a specific source at deployment without retraining, while preserving shared knowledge and general language capabilities.
Read source
Your take?
PapersAI safetyAlignmentFine-tuning

Summary generated by Claude — human-verified