Retour au feed
Reddit r/LocalLLaMA·

Qwen3.6-27B on 2x3090s: llama.cpp vs vLLM, all the flags, and the MTP acceptance/inference speed/context

Signal
78
Hype
15
En 3 lignesBenchmark détaillé de Qwen3.6-27B sur 2x RTX 3090 comparant llama.cpp (Q6_K/Q8_0) et vLLM (INT4/INT8). Mesures réelles : débit 43-54 tok/s, taux d'acceptation MTP 27-77% selon backend. Setup avec proxy OpenAI-compatible hot-swappant 4 configurations, sans P2P PCIe (Threadripper 1950X).
Lire la source
Ton avis ?
QwenGénération de codeBenchmarksOutilsOpen source

Résumé généré par Claude — vérifié par l'humain