Back to feed
arXiv cs.CL·

Learning What to Learn: Stage-Specific Data Sets for SFT-then-RL in Small Language Model Reasoning

Signal
75
Hype
15
In three linesSFT-then-RL training framework for small language models: SFT acquires not-yet-mastered reasoning skills, RL consolidates them. Bridge mechanism transforms raw reasoning traces into learnable supervision. Critique Fine-Tuning converts zero-reward failures into diagnostic supervision. Consistent improvements across five reasoning benchmarks.
Read source
Your take?
Fine-tuningReinforcement learningReasoningBenchmarks

Summary generated by Claude — human-verified