acceptodds
Under review as a conference paper at ICLR 2027

The Best Optimizer Depends on Batch Size

Abstract

A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning. To understand these failures, we analyze a noisy quadratic model and show that the critical batch size beyond which a scaling rule can break down varies across directions in the loss landscape. We further prove that gradient noise drives a bias–variance tradeoff in preconditioning that can reverse optimizer rankings across batch sizes, even under optimal learning rate and momentum tuning. Taken together, our results show that principled optimizer design must account for batch size and empirical benchmarking should be performed across batch sizes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.