PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation
Signal
78
Hype
15
In three linesPowerOPD stabilizes on-policy distillation for LLMs by replacing unbounded log-ratio rewards with Box-Cox power transformation. On 6 mathematical reasoning benchmarks with Qwen3, achieves +6.37 Avg@8/+5.71 Pass@8 gains vs vanilla OPD, reduces wall-clock time by 59.2% and peak GPU memory by 23.1%.Read source
Your take?
Summary generated by Claude — human-verified