Muon–Adam Training for Overparameterized Neural Networks: Convergence and Generalization
Abstract
Muon, an optimizer designed for matrix parameters, has recently demonstrated strong empirical performance in large language model pretraining. However, Muon's underlying mechanisms remains elusive, particularly in practical settings that combine it with other optimizers. In this paper, we study the convergence and generalization of joint Muon–Adam training for overparameterized ReLU networks with two layers. Specifically, we apply Muon to the hidden weight matrix and Adam to the output weights. We show that, with paired random initialization, the empirical loss converges exponentially to zero. Our analysis reveals and separately characterizes the contributions of Muon and Adam to convergence. Furthermore, we establish a generalization bound throughout training, which is determined by the kernel complexity of the training data and reduces to after sufficiently many iterations when this complexity is uniformly bounded. Our results shed light on the roles of Muon and Adam in joint neural network training from a theoretical perspective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.