Back to feed
Reddit r/LocalLLaMA·

Gemma 4 QAT Q4_0 Bench on Strix Halo

Signal
72
Hype
15
In three linesGemma 4 QAT Q4_0 benchmark on Strix Halo APU via llama.cpp Vulkan/RADV. Models tested: 12B (6.50 GiB), 26B-A4B (13.45 GiB), 31B (16.44 GiB). QAT preserves original model behavior better than post-training quantization. QAT assistant heads converted to GGUF for improved acceptance.
Read source
Your take?
GeminiOpen sourceBenchmarksInfrastructure

Summary generated by Claude — human-verified