Learning What to Learn: Stage-Specific Data Sets for SFT-then-RL in Small Language Model Reasoning
Signal
75
Hype
15
In three linesSFT-then-RL training framework for small language models: SFT acquires not-yet-mastered reasoning skills, RL consolidates them. Bridge mechanism transforms raw reasoning traces into learnable supervision. Critique Fine-Tuning converts zero-reward failures into diagnostic supervision. Consistent improvements across five reasoning benchmarks.Read source
Your take?
Summary generated by Claude — human-verified