Optimizer-dependent training dynamics converge to the same one-third optimal data scaling
Abstract
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size . We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to across optimizers. In an online teacher–student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents and . Under SGD, both are close to , so the data exponent is also across different learning rates. Under Adam the two separate: but . Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: -independent for SGD but falls with for Adam. Yet tuned to that optimum, the loss returns to for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, , which fixes the optimal data exponent at . Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.