acceptodds
Under review as a conference paper at ICLR 2027

Gradient Cluster Imbalance: Why Large Batches Favor Normalized Optimizers

Abstract

Adam and Muon often outperform plain stochastic gradient descent in neural network training, but their advantages vary across layers and batch sizes. Previous work has found that Adam is robust to class imbalance in the first and last layers, while Muon is effective in hidden layers. However, a unifying mechanism remains unclear, particularly how these advantages depend on batch size. We identify a property for per-sample gradients that we refer to as gradient cluster imbalance: per-sample gradients form clusters with similar directions that are nearly orthogonal to other clusters, and those clusters have imbalanced sizes. Under this property, we show that GD makes slow progress in directions corresponding to rare clusters for smooth, convex objectives, and that large batches offer limited gains over stochastic gradient descent as averaging leads to ill-conditioning. Under additional structural assumptions, SignGD, a simplified model of Adam, and idealized Muon mitigate this imbalance by normalizing gradients from different clusters separately. Measurements of individual layers provide partial support for these structural assumptions. Experiments reveal unequal gradient clusters across hidden layers of language models and vision models, extending existing observations of frequency imbalance beyond class labels. We show that clusters vary across layers and over training and reflect shared token patterns and visual features.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.