Back to feed
arXiv cs.CL·

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

Signal
72
Hype
28
In three linesSHARD is a self-reframing distillation method to improve safe-helpfulness balance in LLMs. It rewrites sensitive prompts using philosophical guidelines to surface benign intent, reframes responses into safer and more helpful versions, then fine-tunes the model on self-reframed responses. Tested on DNA and LINGUASAFE, SHARD improves helpfulness while preserving safety.
Read source
Your take?
Fine-tuningAI safetyAlignmentPrompt engineering

Summary generated by Claude — human-verified