Accelerating Muon with Spectral Residual Correction
Abstract
Muon constructs matrix updates by applying a matrix-sign transformation to momentum. A common route to accelerating Muon is to modify the gradient or momentum before this transformation. We explore a complementary direction: refining the parameter step while retaining Muon's momentum recursion. Derived from the parameter subproblem of the ADMM-Inspired Momentum framework, AIMuon uses the gradient–momentum residual to correct the tentative spectral step. The residual undergoes a separate matrix-sign transformation, while spectral–nuclear duality determines the relative weighting of the momentum and residual directions. In practice, both directions are approximated with Newton–Schulz iterations and their combination is rescaled to match Muon's update magnitude. We establish a stochastic stationarity bound for the corresponding exact iteration. On GPT-2-small and GPT-2-medium pretraining, AIMuon reaches Muon's final validation loss with 12.8% and 7.7% fewer optimization steps, respectively, while achieving lower final validation loss than both Muon and MONA. Ablation studies further support the roles of the residual correction, separate transformation, adaptive nuclear-norm weighting, and final normalization. These results show that parameter-step refinement can complement momentum construction in matrix-sign optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.