Spectral Flattening Is Not Enough: When Muon Improves Learning Rates and Convergence
Abstract
Muon updates matrix-valued parameters by flattening the singular values of each gradient while preserving its singular directions. Despite its practical success, the mechanisms governing its learning-rate capacity and convergence remain insufficiently understood. We analyze deterministic, full-batch Muon through its exact polar update. First, we derive a descent-preserving learning-rate threshold and show that it is controlled by the block-size-weighted average singular scale of the gradient matrices. This characterization makes the role of gradient magnitude explicit and reveals how objective scaling affects the admissible step size. We then analyze Muon's convergence behavior and identify the spectral properties that govern its convergence rate. Our analysis shows that flattening the spectrum within individual parameter blocks is not sufficient to determine global convergence, because scale differences across blocks remain influential. Together, these results provide a theoretical account of how spectral flattening governs the learning-rate capacity and convergence behavior of Muon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.