Be wary of Qwen/Claude distillations - they're often worse than the base model
Signal
72
Hype
45
In three linesQwen/Claude distillations circulating on r/LocalLLaMA (Qwopus, Fable 5 on Qwen 3.6) use 4k-10k training samples, insufficient to improve performance. Compared to 700k samples in official DeepSeek-R1 distillations, these models don't exceed base Qwen and slightly degrade quality despite different reasoning style.Read source
Your take?
Summary generated by Claude — human-verified