Residual-Augmented Scaling Laws for LLM Pretraining Performance Prediction
Abstract
The computational cost of pretraining Large Language Models (LLMs) at modern scales makes it essential to predict pretraining outcomes (validation loss on hold-out data) before committing to expensive full-scale pretraining. The scaling laws of parameters and token sizes enable extrapolation from small-scale pilots but miss other pretraining feature-dependent variations (i.e., hyperparameters and data states), while purely data-driven methods require substantial training data and extrapolate unreliably beyond observed regimes. To address this dual limitation, we propose ReScaL , a hybrid method that diversely fits an ensemble of scaling laws on minimal pilot runs; learns ML-predicted residual corrections to capture pretraining feature-dependent variations that scaling laws miss; and dynamically combines scaling law extrapolation with residual corrections to leverage the strengths of both. Through experiments on 52 pairs of pretraining datasets and budget sizes, we show that ReScaL is ranked 1st in 41/52 cases (79%) with the best average rank of 1.21, and attains up to 5.7 lower MAPE than the strongest counterpart. Ablations show that removing each design can degrade accuracy by up to 4.3, while sensitivity analysis confirms our hyperparameter choices as robust operating points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.