Back to feed
Reddit r/LocalLLaMA·

Jetbrains Mellum 2: a really good and performant model

Signal
65
Hype
35
In three linesJetBrains Mellum 2, a 12B MoE model with 2.5B active params, achieves 111.2 t/s generation and 492.7 t/s prompt eval on AMD RX 7900 XT. Outperforms Gemma 4-12B and GPT-OSS-20B on complex tool-calling tasks. 3.7× faster than Qwen 3.5-9B on same hardware.
Read source
Your take?
Open sourceBenchmarksCode generationAI Agents

Summary generated by Claude — human-verified