Spokes: Optimizing for Diverse Pretraining Data Selection
Signal
78
Hype
15
In three linesSPOKES optimizes pretraining data selection through a probabilistic diversification framework based on G-Vendi score and exponentiated gradient descent. On FineWeb and DCLM, the method improves downstream performance by +1.5 and +1.4 points when jointly optimizing quality and diversity, outperforming semantic deduplication.Read source
Your take?
Summary generated by Claude — human-verified