acceptodds
Under review as a conference paper at ICLR 2027

From Data Imbalance to Optimizer Differences: A Hessian Perspective

Abstract

Prior work has used data imbalance to explain performance differences between optimizers, but frequency statistics alone are insufficient to characterize the optimization difficulties faced by different optimizers. In this work, we connect data imbalance to the blockwise Hessian structure that shapes optimizer behavior. Specifically, we characterize this structure through two properties—block-scale heterogeneity and within-block anisotropy—and find that optimizers respond differently to them: Adam cancels block-scale factors but does not necessarily alleviate within-block anisotropy, whereas Muon can somewhat mitigate the optimization difficulties arising from both properties. To understand how these two properties arise, we study a two-layer linear bigram model. Under mean-squared error (MSE), the Hessians depend on the data distribution only through input frequencies: input imbalance induces block-scale heterogeneity in embedding training, while shaping within-block anisotropy in output-head training. Under cross-entropy (CE), output frequencies also play a role in shaping these properties. Together, these results show that data imbalance does not act alone: it interacts with the choice of trainable parameters and loss function to shape these two Hessian properties, thereby influencing the relative performance of different optimizers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.