Rethinking Pre-pretraining
Abstract
We report that allocating compute to synthetic pre-pretraining (PPT) never outperforms pure natural-language training in compute-matched experiments, across four PPT curricula (-Dyck, neural cellular automata (NCA), declarative reference corpora, and code) and three training regimes: data-constrained, data-unconstrained, and a high token-to-parameter-ratio regime (881M parameters trained on 50B tokens). When natural language is scarce, repeated epochs over a 100M-token corpus beat every synthetic curriculum on natural-language validation loss, with no consistent downstream accuracy advantage from PPT. We find that broader PPT gains reported in prior work can be attributed to the weight initialization scheme. In an NCA reproduction, we recover a 5.4% PPT advantage under the original initialization; changing the initialization improves both methods (the no-PPT baseline by 9.0%), leaving them indistinguishable. However, when PPT compute is *additive* to, rather than *competing* with, the natural-language budget, PPT gains persist even after improving the baseline initialization at small natural-language budgets, but vanish by 1B natural-language tokens. Finally, knowledge acquired during PPT fails to survive subsequent natural-language training; interleaving the same data throughout training preserves it, though it does not significantly improve performance. Our results suggest limited benefit from synthetic pre-pretraining as a distinct training stage and motivate rethinking how such synthetic data are used in language model training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.