Preconditioning Benefits of Spectral Orthogonalization in Muon
Abstract
The Muon optimizer has achieved strong empirical performance in large language model training, yet the role of spectral orthogonalization in its effectiveness remains poorly understood. We study its preconditioning effect through two case studies: matrix factorization and in-context linear regression in an infinite-prompt Gaussian softmax-attention model. For both problems, we prove that exact-polar \muon attains logarithmic iteration complexity independent of the relevant condition number, for every fixed momentum parameter under suitable initialization and stepsize conditions. We also establish condition-number-dependent lower bounds for gradient descent and standard full-batch Adam on separate hard instances. On a common locally initialized matrix-factorization instance, we further separate Muon's logarithmic dependence on the condition number from stabilized Adam's square-root dependence, assuming bounded scalar learning rates. Our analysis reduces the convergence argument to scalar dynamics in the spectral domain, revealing how orthogonalization balances progress across modes. These results identify spectral orthogonalization as the source of the preconditioning effect, which arises even without momentum.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.