Best Before Post-Training? The Shelf Life of Pretraining Mixture Rankings
Abstract
The pretraining data mixture strongly influences a language model's capabilities, yet prior research on data mixture optimization typically selects candidates before observing their post-training performance. This practice rests on the untested assumption that pretraining rankings survive post-training. In this paper, we conduct a thorough experimental study over 106 models of 1.5B parameters that differ only in their pretraining mixtures and share identical mid-training and supervised fine-tuning (SFT) data and configurations. Our results demonstrate that knowledge rankings change across these stages, while math rankings largely persist but are diluted by a fixed-weight composite of pretraining scores. Yet these scores retain useful information: pretraining GSM8K alone predicts post-SFT rankings better than the composite for knowledge, math, and code. We therefore introduce MixSelector, a ridge-regression method that uses a subset of post-trained candidates to learn how to weight the pretraining scores together with the mixture composition and then ranks the remaining candidates. On held-out candidates, its ranking agrees with the post-SFT ranking at a mean Spearman correlation of 0.72 across capability families, against 0.10 for the composite, while maintaining a smaller lead over the best single score. With only ten post-trained candidates, it retains most of this advantage. Fixed pretraining rankings thus have a limited shelf life, but recalibrating pretraining signals on a few post-trained models suffices to align mixture selection with post-training performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.