Back to feed
arXiv cs.LG·

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier

Signal
78
Hype
15
In three linesPROPEL is a framework training task generators via RL to create optimally difficult problems for agent learning. A lightweight probe predicts solver pass rate without repeated rollouts, reducing evaluation to a single forward pass. On code and SWE tasks, learnable-frontier generation increases from 10.1% to 20% (Qwen2.5-3B) and 9.8% to 19.6% (Qwen3.5-27B).
Read source
Your take?
Reinforcement learningAI AgentsCode generationBenchmarksQwen

Summary generated by Claude — human-verified