Back to feed
Reddit r/LocalLLaMA·

Qwen3.6-27B on 2x3090s: llama.cpp vs vLLM, all the flags, and the MTP acceptance/inference speed/context

Signal
78
Hype
15
In three linesDetailed benchmark of Qwen3.6-27B on 2x RTX 3090 comparing llama.cpp (Q6_K/Q8_0) and vLLM (INT4/INT8). Real measurements: throughput 43-54 tok/s, MTP acceptance rates 27-77% per backend. Setup with OpenAI-compatible proxy hot-swapping 4 configs, no PCIe P2P (Threadripper 1950X).
Read source
Your take?
QwenCode generationBenchmarksToolsOpen source

Summary generated by Claude — human-verified