On the convergence analysis of the AdaMuon optimizer
Abstract
In recent years, Muon has demonstrated excellent experimental performance in neural networks. This success has given rise to numerous Muon variants, among which AdaMuon introduces a dynamic adaptive mechanism for large-scale neural network training. However, despite its strong empirical performance, the convergence guarantees hold only under some specific assumptions and are not fully established. To to further narrow the theoretical gap, this paper provides a systematic convergence analysis of AdaMuon. Specifically, we establish convergence analysis under three distinct settings. For nonconvex case under and without Frobenius norm Lipschitz smoothness, we prove the convergence rate is \(O(T^-1/4)\). For star-convex case under Frobenius norm Lipschitz smoothness, we prove the convergence rate to reach the precision \(\epsilon\) is \(O(\epsilon^-3\log(\epsilon^-1))\). These results extend the theoretical scope of the original analysis, offering a more complete characterization of AdaMuon across different optimization assumptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.