Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
Signal
78
Hype
25
In three linesarXiv paper on improving long-context reasoning via data-centric approach rather than reward engineering. Data recipe targeting retrieval, multi-evidence synthesis, reasoning (~14K examples). Tests on Qwen3 (4B/8B/30B): +7.2/+3.2/+6.4 points across 7 long-context benchmarks, transfer to agentic tasks (+4.8 GAIA, +7.0 BrowseComp).Read source
Your take?
Summary generated by Claude — human-verified