Back to feed
arXiv cs.LG·

PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation

Signal
78
Hype
15
In three linesPowerOPD stabilizes on-policy distillation for LLMs by replacing unbounded log-ratio rewards with Box-Cox power transformation. On 6 mathematical reasoning benchmarks with Qwen3, achieves +6.37 Avg@8/+5.71 Pass@8 gains vs vanilla OPD, reduces wall-clock time by 59.2% and peak GPU memory by 23.1%.
Read source
Your take?
Fine-tuningReinforcement learningBenchmarksPapers

Summary generated by Claude — human-verified