Back to feed
Reddit r/LocalLLaMA·

2X tk/s (from 19.4 -> 38.1 tk/s on 1 x MI50) Playing with a hypothesis like speculative decoding.. but instead of an additional side model, exploiting that I can run multiple computations side-by-side AS IF I had Qwen3.6-27B loaded twice in memory - small quants don't use all the available compute.

Signal
65
Hype
25
In three linesSpeculative decoding optimization: 2x throughput increase (19.4 → 38.1 tk/s on MI50) by running multiple parallel computations with the same Q8 quantized model, exploiting that small quantizations use only 25% of available compute.
Read source
Your take?
Open source

Summary generated by Claude — human-verified