Sharpness-Aware Muon: Matrix-Aware Optimization with Momentum-Guided Perturbations
Abstract
Muon is a matrix-aware optimizer that combines momentum with orthogonalization to produce structured update directions with strong empirical behavior. However, it only shapes the geometry of the update itself and does not explicitly control neighborhood worst-case loss or bias training toward flat regions, which are often linked to stronger generalization. To address this limitation, we introduce the sharpness-aware formulation of Muon. We first develop MuSAM, which incorporates Sharpness-Aware Minimization (SAM)-style perturbation into Muon's orthogonalized update and establishes a principled bridge between sharpness-aware training and matrix-aware optimization. We then show that this formulation faces an efficiency–stability tension: constructing the perturbation requires an extra gradient evaluation, while the perturbation itself is generated from a single minibatch gradient and is therefore noisy. To resolve this issue, we propose MoMuSAM, which turns Muon's momentum from an update variable into a sharpness-probing signal. By using normalized momentum as the perturbation direction, MoMuSAM removes the extra clean-gradient computation and replaces the single-batch perturbation proxy with a temporally aggregated one. We prove that both MuSAM and MoMuSAM attain the same asymptotic first-order stationarity order, and further show that MoMuSAM yields a more stable stochastic linearization than MuSAM. Experiments on standard benchmarks demonstrate that sharpness awareness consistently improves Muon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.