On the Convergence and Divergence of Muon: The Role of Momentum
Abstract
Muon has emerged as an effective optimizer for training large language models. Previous studies have established momentum-dependent convergence guarantees for Muon. Some obtain convergence by choosing approaching one as the training horizon grows. Other analyses give guarantees for arbitrary fixed , but often require growing batch sizes or additional conditions such as noise that vanishes with the gradient norm. These results do not establish both a convergence threshold independent of the training horizon and small-momentum divergence on the same fixed problem. To our knowledge, we provide the first joint analysis of these two behaviors for Muon within a fixed smooth finite-sum setting. We establish sufficient high-momentum thresholds and for best-iterate stationarity bounds under the affine growth condition and diminishing stepsizes: pathwise for arbitrary epochwise permutations and in expectation for i.i.d. uniform sampling. Under the strong growth condition, these bounds tend to zero as the horizon grows. We also establish a sufficient low-momentum threshold and a smooth, strongly convex quadratic for which, from suitable initializations, Muon diverges to infinity for all : on every permutation path and on an event of arbitrarily high probability under i.i.d. sampling. Within each sampling scheme, the high-momentum guarantees apply to the same quadratic with the initialization and stepsize schedule unchanged. Together, these results show that changing momentum alone can separate triple divergence from vanishing stationarity bounds on the same fixed problem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.