SHARD: Safe and Helpful Alignment via Self-Reframing Distillation
Signal
72
Hype
28
In three linesSHARD is a self-reframing distillation method to improve safe-helpfulness balance in LLMs. It rewrites sensitive prompts using philosophical guidelines to surface benign intent, reframes responses into safer and more helpful versions, then fine-tunes the model on self-reframed responses. Tested on DNA and LINGUASAFE, SHARD improves helpfulness while preserving safety.Read source
Your take?
Summary generated by Claude — human-verified