K-Scaling: Critical Batch Size Governs Learning Rate Scheduling Across Token Budgets in LLM Pretraining
Abstract
We study learning rate (LR) scheduling in large language model pretraining by jointly tuning peak LR and decay ratio, the fraction of post-warmup training allocated to terminal decay. Using the functional scaling law framework with an admissibility constraint on peak LR, we establish : when expressed through the normalized batch size (BS) , both optimal hyperparameters follow budget-independent scaling profiles, where is the token budget and is a shared critical batch size (CBS). The CBS marks a from pure decay with increasing peak LR in the regime () to warmup–stable–decay (WSD) with saturated peak LR and decreasing decay ratio in the regime (). Experiments on dense and MoE models spanning 0.1B–1B parameters, trained with MuonH, confirm approximate K-scaling collapse and support a practical four-parameter model for jointly predicting peak LR and decay ratio. Fitted on small token budgets, the model extrapolates to budgets up to the largest fitting budget, yielding schedules with final losses close to grid-search optima without further tuning. The theory also explains why longer training favors larger decay ratios at fixed-BS training, helping reconcile differing schedule preferences in prior studies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.