Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
Abstract
Neural scaling laws are foundational for language model development, yet the standard formulation systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. We show that this interaction is directly measurable in the loss surface, and capture it with the Skaling law (/ska.liN/}), a coupled functional form that extends Chinchilla with a single interaction exponent between model capacity and data. On two large training grids, this one extra parameter reduces the Mean Absolute Percentage Error (MAPE) by roughly – on interpolation and extrapolation, and by up to an order of magnitude when fitted on sparse grids. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law matches or exceeds the accuracy of Chinchilla fitted on the full grid while using up to less compute. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.