acceptodds
Under review as a conference paper at ICLR 2027

Best Before Post-Training? The Shelf Life of Pretraining Mixture Rankings

Abstract

The pretraining data mixture strongly influences a language model's capabilities, yet prior research on data mixture optimization typically selects candidates before observing their post-training performance. This practice rests on the untested assumption that pretraining rankings survive post-training. In this paper, we conduct a thorough experimental study over 106 models of 1.5B parameters that differ only in their pretraining mixtures and share identical mid-training and supervised fine-tuning (SFT) data and configurations. Our results demonstrate that knowledge rankings change across these stages, while math rankings largely persist but are diluted by a fixed-weight composite of pretraining scores. Yet these scores retain useful information: pretraining GSM8K alone predicts post-SFT rankings better than the composite for knowledge, math, and code. We therefore introduce MixSelector, a ridge-regression method that uses a subset of post-trained candidates to learn how to weight the pretraining scores together with the mixture composition and then ranks the remaining candidates. On held-out candidates, its ranking agrees with the post-SFT ranking at a mean Spearman correlation of 0.72 across capability families, against 0.10 for the composite, while maintaining a smaller lead over the best single score. With only ten post-trained candidates, it retains most of this advantage. Fixed pretraining rankings thus have a limited shelf life, but recalibrating pretraining signals on a few post-trained models suffices to align mixture selection with post-training performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.