Data–Task Alignment as a Third Scaling Axis for Language-Model Fine-Tuning
Abstract
Scaling laws for language models usually predict loss from model size and training tokens while treating the data distribution as fixed. This assumption is especially limiting in fine-tuning, where examples at the same token budget can differ substantially in their relevance to the target task. Existing laws cannot quantify that difference. We introduce data–task alignment , a measurable property of a training set relative to a held-out target, and study it jointly with parameter count and fine-tuning tokens . Our primary experiment contains 51 controlled Qwen3 cells across four model sizes, six alignment conditions, and a token range. Within this grid, the first-order law predicts held-out loss with and RMSE . Omitting alignment changes both fitted compute exponents, so heterogeneous data composition can be misattributed to scaling with and . We then map the law's operating regime through a crossed-budget study that finds mild dependence on the token-to-parameter ratio. Selection beyond Q90 shows diminishing returns and eventual reversal as the eligible pool narrows. The alignment effect replicates on an independent math target, a code domain, Llama-3.2, and a frozen gradient selector, although the preferred functional form changes. Alignment therefore provides a useful third scaling coordinate whose interpretation remains tied to the target and measurement. These results connect task-specific data selection to finite-budget scaling by showing when better-targeted data can substitute for compute. They also identify how this return changes across regimes and eventually reverses when selection becomes too narrow.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.