Back to feed
arXiv cs.LG·

Temporal Difference Learning for Diffusion Models

Signal
72
Hype
18
In three linesNovel training approach for diffusion models using temporal difference (TD) objective to enforce multi-step consistency along the denoising trajectory. Reformulates diffusion as a Markov reward process and denoising as policy evaluation in reinforcement learning. Shows significant FID improvements, especially with few sampling steps.
Read source
Your take?
Reinforcement learningReasoning

Summary generated by Claude — human-verified