acceptodds
Under review as a conference paper at ICLR 2027

Preconditioning Benefits of Spectral Orthogonalization in Muon

Abstract

The Muon optimizer has achieved strong empirical performance in large language model training, yet the role of spectral orthogonalization in its effectiveness remains poorly understood. We study its preconditioning effect through two case studies: matrix factorization and in-context linear regression in an infinite-prompt Gaussian softmax-attention model. For both problems, we prove that exact-polar \muon attains logarithmic iteration complexity independent of the relevant condition number, for every fixed momentum parameter under suitable initialization and stepsize conditions. We also establish condition-number-dependent lower bounds for gradient descent and standard full-batch Adam on separate hard instances. On a common locally initialized matrix-factorization instance, we further separate Muon's logarithmic dependence on the condition number from stabilized Adam's square-root dependence, assuming bounded scalar learning rates. Our analysis reduces the convergence argument to scalar dynamics in the spectral domain, revealing how orthogonalization balances progress across modes. These results identify spectral orthogonalization as the source of the preconditioning effect, which arises even without momentum.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.