Back to feed
Reddit r/LocalLLaMA·

Gemma 4 models benchmarked on with Triple GPU

Signal
72
Hype
15
In three linesGemma 4 benchmarked on triple GPU setup (3× GTX-1070, 24 GiB VRAM total). Gemma-4-26B-A4B-qat achieves 123.5 t/s prompt processing and 53.08 t/s generation. Gemma-4-E4B-BF16 reaches 302.16 t/s but limited to 11.54 t/s generation. Tests on llama.cpp build 9204 with GGUF quantizations.
Read source
Your take?
GeminiBenchmarksOpen sourceInfrastructure

Summary generated by Claude — human-verified