Back to feed
arXiv cs.LG·

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR

Signal
78
Hype
15
In three linesStudy shows SFT overtraining can invert model rankings during RLVR fine-tuning. On Qwen2.5-Coder-3B, increasing SFT depth raises pre-RL pass@1 but reduces GRPO pass@10 from 0.806 to 0.481. Pre-RL entropy positively correlates with RLVR outcomes (ρ=+0.69). Two-stage entropy-based diagnostic identifies high-risk checkpoints.
Read source
Your take?
Reinforcement learningFine-tuningReasoningBenchmarks

Summary generated by Claude — human-verified