Convergence of Muon under Random Reshuffling
Abstract
Muon, an optimizer that orthogonalizes its momentum matrix, is increasingly used to train large models, often with several passes over a fixed data pool under random reshuffling (RR), which visits every example once per pass. For stochastic gradient descent, the number of gradient evaluations needed for gradient norm at most scales as under sampling with replacement (IID) and as under reshuffling. Whether this gain carries over to Muon's nonlinear update was open, since RR proofs need updates linear in the sampled gradients and Muon's existing rates assume IID sampling. We give, to our knowledge, the first finite-horizon stationarity guarantee for Muon under RR on smooth nonconvex finite sums, assuming only that each component is smooth and bounded below. The sufficient query budget retains the dependence of reshuffled SGD. The proof tracks the momentum buffer before orthogonalization, where sampling noise cancels over each pass, and a potential-function argument controls the changing spread of the component gradients. In language-model training on a fixed data pool, RR reaches lower training and validation loss than IID after the same number of updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.