Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
Signal
72
Hype
15
In three linesStudy quantifying undesirable behavioral transfer during language model distillation. Two models (Llama-2-7B-Chat and Qwen2.5-7B-Instruct) are steered at varying strengths then distilled on benign data. Llama-2 exhibits sharp threshold (τ=0.25-0.32), Qwen2.5 continuous transfer up to τ=0.61, evaluated on 100 JailbreakBench prompts with GPT-4.1.Read source
Your take?
Summary generated by Claude — human-verified