Back to feed
Reddit r/LocalLLaMA·

120 tok/s on 12GB VRAM with Gemma 4 12B QAT MTP

Signal
75
Hype
25
In three linesGoogle released Gemma 4 12B QAT (Quantization-Aware Training) variant. User achieves 120 tok/s on RTX 4070 Super 12GB with llama.cpp using MTP speculative decoding and draft assistant model. Detailed benchmark across 9 tasks with 65.78% aggregate acceptance rate.
Read source
Your take?
GeminiCode generationBenchmarksOpen sourceTools

Summary generated by Claude — human-verified