acceptodds
Under review as a conference paper at ICLR 2027

Provable Advantages of Muon from Data-Induced Curvature

Abstract

Muon, a matrix-orthogonalization-based optimizer, has shown strong empirical performance in large-scale language-model training, yet when and why spectral normalization yields a sustained optimization advantage remain incompletely understood. Since Muon acts on matrix-valued parameter blocks, we study a one-hidden-layer quadratic network with a single trainable matrix, which retains the matrix-gradient structure while linking input moments explicitly to parameter curvature. Motivated by empirical feature spectra, we consider Gaussian inputs with geometrically decaying centered variances and a nonzero mean whose direction is close to an active feature. In this model, data geometry induces anisotropic curvature and an imbalanced gradient spectrum, creating a regime favorable to spectral normalization. Under explicit conditions, we prove that exact-polar Muon achieves an optimization advantage over both gradient descent and SignGD, a special case of Adam's coordinatewise normalization, and that this advantage persists along the nonlinear trajectories on a fixed empirical objective. These results establish a provable mechanism linking data-induced curvature to the optimization advantage of spectral normalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.