acceptodds
Under review as a conference paper at ICLR 2027

High-Confidence Stationarity of Momentum Muon: A Heavy-Tailed Separation from Adam

Abstract

As an emerging optimization method, Muon has been empirically shown to outperform Adam/AdamW on modern machine learning tasks, yet the mechanism underlying this advantage remains unclear. We identify a key mechanism behind Muon's high-probability advantage: its instantaneous singular-value compression provides the scale regularization sought by Adam's second-moment normalization while directly controlling unpreconditioned stationarity. By avoiding the confidence loss incurred when converting Adam's preconditioned adaptive energy into unpreconditioned stationarity, Muon achieves sharper confidence dependence under heavy-tailed noise. Specifically, under conditional -moment noise for any , we prove a finite-time high-probability theorem for exact canonical pre-polar momentum Muon, with fixed within each run. Choosing momentum and stepsize for the given horizon at fixed batch size yields the average-gradient rate , where . Even when stochastic gradients have infinite variance, the dependence on the confidence parameter remains only logarithmic. For the five-step Newton–Schulz (NS5) variant, we obtain the same rate up to an additive error floor under an explicit approximation-error condition. In contrast, for Adam with fixed , and initial moments, we construct a uniformly -moment-bounded two-dimensional family and prove a minimax stationarity lower bound with probability greater than .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.