Hyperparameter Transfer and Ensembling Improve Data-Constrained Scaling Recipes
Abstract
As growing compute budgets outpace the growth of high-quality training data, large language models must be trained with both compute and data constraints in mind. Unfortunately, existing prescriptions for data-constrained scaling suffer from two important limitations. First, whereas optimal hyperparameters vary wildly with model size and number of epochs, past works hold key hyperparameters such as weight decay fixed with scale. Second, past works neglect ensembling, a promising axis for scaling under data constraints. We address both limitations, leading to improved prescriptions for scaling data-constrained training pipelines. Our method prescribes how to scale key hyperparameters and allocate compute to scaling epochs, model parameters, or ensemble members. Controlling for training corpus size and compute budgets, our prescriptions achieve significantly lower validation loss than existing baselines, and this gap grows at larger scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.