Back to feed
Reddit r/LocalLLaMA·

Be wary of Qwen/Claude distillations - they're often worse than the base model

Signal
72
Hype
45
In three linesQwen/Claude distillations circulating on r/LocalLLaMA (Qwopus, Fable 5 on Qwen 3.6) use 4k-10k training samples, insufficient to improve performance. Compared to 700k samples in official DeepSeek-R1 distillations, these models don't exceed base Qwen and slightly degrade quality despite different reasoning style.
Read source
Your take?
QwenClaudeFine-tuningBenchmarks

Summary generated by Claude — human-verified