Back to feed
Reddit r/LocalLLaMA·

Scaling former VibeThinker-1.5B to 3B — now it reaches frontier math & coding performance

Signal
75
Hype
35
In three linesVibeThinker-3B achieves 94.3 on AIME'26, 80.2 on LiveCodeBench v6, and 96.1% pass rate on unseen LeetCode contests. The model demonstrates small models can reach frontier-level reasoning performance in math and coding through clear verification signals.
Read source
Your take?
ReasoningBenchmarksCode generationOpen source

Summary generated by Claude — human-verified