acceptodds
Under review as a conference paper at ICLR 2027

Accelerating Muon with Spectral Residual Correction

Abstract

Muon constructs matrix updates by applying a matrix-sign transformation to momentum. A common route to accelerating Muon is to modify the gradient or momentum before this transformation. We explore a complementary direction: refining the parameter step while retaining Muon's momentum recursion. Derived from the parameter subproblem of the ADMM-Inspired Momentum framework, AIMuon uses the gradient–momentum residual to correct the tentative spectral step. The residual undergoes a separate matrix-sign transformation, while spectral–nuclear duality determines the relative weighting of the momentum and residual directions. In practice, both directions are approximated with Newton–Schulz iterations and their combination is rescaled to match Muon's update magnitude. We establish a stochastic stationarity bound for the corresponding exact iteration. On GPT-2-small and GPT-2-medium pretraining, AIMuon reaches Muon's final validation loss with 12.8% and 7.7% fewer optimization steps, respectively, while achieving lower final validation loss than both Muon and MONA. Ablation studies further support the roles of the residual correction, separate transformation, adaptive nuclear-norm weighting, and final normalization. These results show that parameter-step refinement can complement momentum construction in matrix-sign optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.