Back to feed
arXiv cs.LG·

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

Signal
72
Hype
15
In three linesStudy quantifying undesirable behavioral transfer during language model distillation. Two models (Llama-2-7B-Chat and Qwen2.5-7B-Instruct) are steered at varying strengths then distilled on benign data. Llama-2 exhibits sharp threshold (τ=0.25-0.32), Qwen2.5 continuous transfer up to τ=0.61, evaluated on 100 JailbreakBench prompts with GPT-4.1.
Read source
Your take?
LlamaQwenFine-tuningAI safetyBenchmarks

Summary generated by Claude — human-verified