Back to feed
arXiv cs.LG·

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Signal
78
Hype
15
In three linesPre-training data mixture experiments fail to scale because repetition rates of high-quality data shift with training budget. A subsampling procedure matching target repetition rates recovers optimal mixtures using only 1/16 of target tokens (757M model), reducing error from 0.75 to 0.05 compared to uncontrolled baselines.
Read source
Your take?
BenchmarksPapers

Summary generated by Claude — human-verified