acceptodds
Under review as a conference paper at ICLR 2027

Muon–Adam Training for Overparameterized Neural Networks: Convergence and Generalization

Abstract

Muon, an optimizer designed for matrix parameters, has recently demonstrated strong empirical performance in large language model pretraining. However, Muon's underlying mechanisms remains elusive, particularly in practical settings that combine it with other optimizers. In this paper, we study the convergence and generalization of joint Muon–Adam training for overparameterized ReLU networks with two layers. Specifically, we apply Muon to the hidden weight matrix and Adam to the output weights. We show that, with paired random initialization, the empirical loss converges exponentially to zero. Our analysis reveals and separately characterizes the contributions of Muon and Adam to convergence. Furthermore, we establish a generalization bound throughout training, which is determined by the kernel complexity of the training data and reduces to after sufficiently many iterations when this complexity is uniformly bounded. Our results shed light on the roles of Muon and Adam in joint neural network training from a theoretical perspective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.