BeeLlama v0.3.1 – latest llama.cpp with extras! DFlash, MTP, q6_0 cache, TurboQuant. Single RTX 3090: Qwen 3.6 27B & Gemma 4 31B up to 177.8 tps (4.93x over baseline)
Signal
72
Hype
35
In three linesBeeLlama v0.3.1 updates llama.cpp with MTP, Gemma 4 12B support, multi-GPU DFlash, and new cache options (q6_0, TQ3_1S, TQ4_1S). On RTX 3090, Qwen 3.6 27B reaches 177.8 tps (4.93x baseline), Gemma 4 31B also optimized. Prebuilt binaries and Docker images provided.Read source
Your take?
Summary generated by Claude — human-verified