HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)
Signal
72
Hype
35
In three linesHalBench v2.3 benchmarks 29 open-source models on sycophancy and hallucination across 3,076 audited questions with false premises. Qwen 3.6 (~27B) scores 36.6% pushback, outperforming all larger open models, GPT-5.4, and Gemini 3.1 Pro. Only Sonnet 4.6 and Grok exceed 50%. Phi-4 scores 2.3%.Read source
Your take?
Summary generated by Claude — human-verified