Generalization Analysis of Muon
Abstract
Muon is a matrix-aware first-order optimizer designed for matrix-structured parameters and has demonstrated strong performance in training large neural networks. Existing analyses mainly study its convergence behavior on training examples, whereas its behavior on testing examples remains less understood. In this paper, we study the convergence of Muon measured by the nuclear norm of the population gradient. We first establish a data-dependent uniform convergence guarantee for the deviation between population and empirical gradients over a ball in with rank at most . Our bound depends on instead of the ambient dimension and incorporates the empirical second-order moment of the gradients. We then establish high-probability convergence rates for Muon. Our generalization and convergence analyses together give a high-probability excess risk bound of order . We further show that the dimension dependency can be removed for linear networks and shallow neural networks. Our studies give the first high-probability generalization guarantee for Muon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.