Read the Trace, Steer the Path: Trajectory-Aware Reinforcement Learning for Diffusion Language Models
Signal
78
Hype
15
In three linesCAPR is a reinforcement learning algorithm for diffusion language models that leverages the denoising trace to generate fine-grained supervision signals without full tree expansion cost. The approach reduces rollout cost to 0.75x flat methods and 0.6x tree methods, achieving SOTA on 4x4 Sudoku, Countdown, GSM8K, and Math500.Read source
Your take?
Summary generated by Claude — human-verified