Retour au feed
Reddit r/LocalLLaMA·

HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)

Signal
72
Hype
35
En 3 lignesHalBench v2.3 évalue 29 modèles open-source sur la sycophantie et hallucinations via 3,076 questions avec fausses prémisses. Qwen 3.6 (~27B) atteint 36.6% de rejet, surpassant tous les modèles open plus grands, GPT-5.4 et Gemini 3.1 Pro. Seuls Sonnet 4.6 et Grok dépassent 50%. Phi-4 obtient 2.3%.
Lire la source
Ton avis ?
BenchmarksOpen sourceÉvaluationsSécurité IAQwen

Résumé généré par Claude — vérifié par l'humain