Implicit Bias of Muon with Finite Newton–Schulz Iterations
Abstract
Muon has recently attracted growing interest as an effective optimizer for large-scale neural network training. However, the theoretical mechanisms underlying Muon, especially the role of Newton–Schulz (NS) iterations in practice, are not yet fully understood. We focus on Muon’s implicit bias with finite NS iterations in multi-class classification. We characterize the gap to the optimal spectral margin through a margin bound, which decomposes the gap into a vanishing training term and an additional path-weighted distortion arising from NS iterations. Persistent distortion can leave a nonvanishing margin gap, whereas vanishing weighted distortion guarantees recovery of the optimal spectral margin. This characterization clarifies when finite NS iterations alter or preserve Muon’s spectral implicit bias, providing a principled understanding of their role in practical Muon implementations. Finally, extensive experiments across synthetic and real-data settings validate our theoretical findings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.