Training-Free Data Mixture Optimization via Neural Tangent Kernels
Abstract
Determining the optimal data mixture is an important problem in continual pre-training and fine-tuning of large language models (LLMs), as it often governs the trade-off between preserving general capabilities and improving domain-specific expertise. Existing approaches usually estimate the optimal mixture through proxy models, scaling laws, or model merging, while directly modeling how the target LLM evolves under different data mixtures remains under-explored. In this paper, we propose a theoretical framework for data mixture optimization that, to our knowledge, the first time directly models the target LLM's own training dynamics. Specifically, we characterize training in the output space and leverage the neural tangent kernel (NTK) to transform discrete gradient updates into continuous logit dynamics. To make the resulting dynamics tractable, we theoretically justify kernel freezing when the NTK direction remains approximately stable throughout training, allowing the time-varying NTK to be approximated by its initial value and reducing the dynamics to a fixed-kernel ordinary differential equation. We then solve the resulting system using an adaptive numerical ODE solver to generate predicted logit trajectories for ranking candidate mixtures without training the target model. Extensive experiments show the effectiveness of our framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.