Back to feed
arXiv cs.CL·

Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation

Signal
78
Hype
25
In three linesNew OPDLM method transforms autoregressive language models into diffusion models without full retraining. Via on-policy distillation, student model generates its own trajectories while frozen original model provides target logits. Result: 15x to 7,000x fewer training tokens required.
Read source
Your take?
Fine-tuningReasoningPapers

Summary generated by Claude — human-verified