AdamW Scaling Laws from First Principles
Abstract
Scaling language-model pre-training to larger token budgets and model sizes requires retuning optimizer hyperparameters, yet existing transfer rules are predominantly empirical and offer limited theoretical guidance. We study this problem through the convergence behavior of SGD with momentum (SGDM) and AdamW under smoothness, bounded stochastic-gradient variance, and the -Kurdyka-Lojasiewicz condition. For both algorithms, we establish convergence guarantees and derive their optimization error under a fixed token budget . In the long-horizon regime, our analysis yields token-budget-aware BST scaling rules: the batch size should grow as , while the learning rate and momentum parameters remain approximately constant. We further prove a matching lower bound for zero-respecting stochastic first-order algorithms, showing that the deterministic and stochastic terms in our upper bounds are unimprovable, demonstrating the tightness of our approach. Experiments on FineWeb with Transformer language models ranging up to 1.4B parameters support these predictions. In particular, after scaling the batch size according to the BST rule, the optimal learning rate transfers across token budgets within each model scale. Together, our results provide a theoretically grounded prescription for transferring hyperparameters to longer pre-training runs without requiring expensive empirical fits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.