Most small-scale software progress in pretraining is data progress
Abstract
How much of the rapid progress in language models observed over the last few years has come from model versus data improvements? Current observational estimates from model releases have only provided measurements of joint improvements from both components. We conduct a novel investigation into the improvements in language model compute efficiency for pretraining specifically from 2019 to 2025, by experimentally decomposing these improvements into algorithmic progress in the training procedure and data progress in the pretraining corpora. We do so by retraining year-representative model training recipes (GPT-2 to OLMo 2) and corpora (OpenWebText to Ultra-FineWeb) at FLOPs. We evaluate these models on a common downstream OLMES10 benchmark suite and compare the observed compute multipliers. From GPT-2 trained on OpenWebText at FLOPs, upgrading the corpus to Ultra-FineWeb yields a multiplier of ×12.0, upgrading the model to OLMo 2 yields a multiplier of ×5.0, and upgrading both ×21.1. A Shapley decomposition attributes 64% of the joint gain to data progress. We conclude that from 2019 to 2025, at the scales we study, most software progress in pretraining is accounted for by data progress.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.