Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws
Signal
78
Hype
15
In three linesStudy of scaling laws for language model pretraining in data-constrained regime. Authors propose MIR (masked-input regularization), an auxiliary next-token prediction loss on randomly masked inputs, and SoftQ, a scaling law coupling model and data size under repeated data. MIR improves validation loss on 72M–1.4B models and equals ~1.3× more unique training data.Read source
Your take?
Summary generated by Claude — human-verified