Back to feed
Reddit r/LocalLLaMA·

Nemotron - King of the Deep? Comparison of 4 models <=120B

Signal
72
Hype
25
In three linesBenchmark of 4 models ≤120B on deep context (up to 400k tokens). Nemotron Super 120B outperforms GPT-OSS 120B and Qwen 3.5 122B in prompt processing (PP) speed from 16-32k tokens onward. Nemotron maintains >100 TPS PP up to 400k context, but token generation (TG) remains slow (10-20 TPS).
Read source
Your take?
BenchmarksQwenOpen source

Summary generated by Claude — human-verified