120 tok/s on 12GB VRAM with Gemma 4 12B QAT MTP
Signal
75
Hype
25
In three linesGoogle released Gemma 4 12B QAT (Quantization-Aware Training) variant. User achieves 120 tok/s on RTX 4070 Super 12GB with llama.cpp using MTP speculative decoding and draft assistant model. Detailed benchmark across 9 tasks with 65.78% aggregate acceptance rate.Read source
Your take?
Summary generated by Claude — human-verified