Provable Checkpoint Reuse for RegMix Proxy Swarms
Abstract
The composition of pre-training data is a primary determinant of language-model quality, which makes mixture optimization a standard stage of large-scale training pipelines. Existing regression-based mixture search methods pioneered by a RegMix-style pipeline fit a surrogate to a swarm of proxy-model runs and therefore train many candidate mixtures from scratch even though their cumulative compositions may overlap heavily. Therefore, we introduce the Mixture Training Reuse (MTR) problem, which aims to realize a fixed family of target data mixtures at minimum training cost by treating every target mixture as a cumulative process to its terminal composition and sharing intermediate checkpoints across runs. For the arbitrary-order problem MTR-a where the order of training examples does not matter, we show that the best shared checkpoint for a set of targets is their coordinatewise minimum, which reduces the continuous checkpoint search to laminar optimization over target subsets; building on this reduction, we prove NP-completeness and a tight -approximation reuse-gain guarantee for a greedy agglomeration algorithm. For the random-order problem MTR-r, we design a descendant-aware greedy planner that admits a checkpoint only when every descendant target satisfies a feasibility test measured by total variation (TV). Our evaluation covers both synthetic mixture families and real-data training. On tractable synthetic instances, MTR-a recovers of the training savings attained by the exact dynamic-programming optimum ( versus ). On a -mixture Essential-Web proxy swarm, MTR-a saves of training tokens, while TV-constrained MTR-r saves with a closer compositional match to independent training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.