Back to feed
Reddit r/LocalLLaMA·

[Benchmark] DFlash Speculative Decoding + KV Cache Compression on RTX 5090 — 3.26x Speedup

Signal
78
Hype
15
In three linesDFlash speculative decoding + KV cache compression benchmark on RTX 5090 with Qwen3.6-27B. 3.26x speedup (turbo4/turbo4), 3.18x (q4_0/turbo4) with only +0.02% PPL degradation. Q5_K_XL outperforms NVFP4-Q8_0. Scripts and raw data available.
Read source
Your take?
QwenBenchmarksOpen source

Summary generated by Claude — human-verified