Small Models Know Which Data Scales
Abstract
Large-scale pre-training requires scalable data selection that can both construct a high-quality training distribution and assess whether its advantage persists at scale. Existing methods typically score documents independently, which gives little control over the distribution of the resulting corpus. They also validate selected data at a few isolated scales, using metrics that are neither comparable across data distributions nor extrapolable across compute. To address these challenges, we propose PAIR (Pairwise Autoregressive Importance Resampling), which uses paired lightweight language models to characterize the Candidate and high-quality Reference corpora and jointly calibrates resampling strength and retention rate to reconstruct the reference distribution under a finite data budget. To assess scaling value, we introduce the Data Scaling Ladder, which evaluates training distributions with benchmark-aligned proxy losses that follow predictable trends across compute and remain comparable across data distributions, enabling lower-compute experiments to forecast target-scale capability. With only a 19M dense scorer, PAIR achieves a 3.9 data-efficiency speedup in continued pre-training of a 30B-A3B MoE model, closely matching the gain obtained with a 1.5B scorer. Across 1e18–1e21 FLOPs, the two selected distributions follow nearly identical scaling trajectories and preserve their advantages on target-aligned capabilities, while lower-compute fits accurately predict the held-out 1e21-FLOP proxy loss. Together, these results show that training data for large-scale models can be effectively modeled, selected, and evaluated at a small fraction of the scale at which they are ultimately consumed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.