acceptodds
Under review as a conference paper at ICLR 2027

When Muon Becomes Gradient Descent: Finite-Time Flows and Noise Smoothing Near the Optimum

Abstract

Muon is a recently proposed optimizer whose matrix update applies the polar factor (matrix sign) of a gradient or momentum matrix. We study an idealized no-momentum, no-weight-decay Muon update on matrix parameters in both deterministic and stochastic regimes. In the deterministic continuous-time limit, Muon is steepest descent with respect to the spectral norm and obeys the dissipation identity ; under a nuclear Łojasiewicz-type error bound this yields *finite-time* convergence, in contrast to the merely exponential rate of gradient flow. In the minibatch case, the correct continuous model is an Itô SDE whose drift is the noise-smoothed polar map , not . For Gaussian noise , the drift linearizes near stationarity as with , and for square matrices. Thus stochasticity smooths the matrix sign and locally turns Muon into a scaled gradient method with an Ornstein–Uhlenbeck noise neighborhood. This suggests a simple hybrid: use Muon away from stationarity to remove anisotropy and switch to SGD near the optimum, where Muon's finite-time advantage disappears and, at a matched local contraction rate, its stationary noise floor is never lower than that of SGD.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.