Topic

#Qwen

Qwen is a family of open-weight language models developed by Alibaba Cloud, spanning text, code, and multimodal tasks. For example, Qwen2.5-72B is a 72-billion-parameter model freely available on Hugging Face.

40Articles
6Sources
60Avg. signal
arXiv cs.LG·

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier

PROPEL is a framework training task generators via RL to create optimally difficult problems for agent learning. A lightweight probe predicts solver pass rate without repeated rollouts, reducing evaluation to a single forward pass. On code and SWE tasks, learnable-frontier generation increases from 10.1% to 20% (Qwen2.5-3B) and 9.8% to 19.6% (Qwen3.5-27B).

Reinforcement learningAI AgentsCode generation
SIG
78
HYP
00
Reddit r/LocalLLaMA·

HalBench: 29 OSS models tested on a custom built Sycophancy and Hallucination Benchmark, Qwen 3.6 and Gemma 4 scoring far above their weight! (While Meta keeps proving they forgot how to spend their money...)

HalBench v2.3 benchmarks 29 open-source models on sycophancy and hallucination across 3,076 audited questions with false premises. Qwen 3.6 (~27B) scores 36.6% pushback, outperforming all larger open models, GPT-5.4, and Gemini 3.1 Pro. Only Sonnet 4.6 and Grok exceed 50%. Phi-4 scores 2.3%.

BenchmarksOpen sourceEvals
SIG
72
HYP
00