Muse: Representation Geometry of Muon Beyond Normalized Momentum
Abstract
Muon applies a spectral transformation to matrix momentum, while normalized SGD with momentum (nSGDM) normalizes its vectorized form. What does the matrix representation contribute beyond vector normalization? We introduce Muse to investigate this question by varying the matrix representation used to compute the update. In 130M and 600M language models, order-preserving reshaping retains lower validation loss than nSGDM. However, fixed-shape entry shuffling removes Muon's advantage over nSGDM in a controlled row-block problem and 600M language-model training. To illustrate how representation affects optimization, we first show that single-step parameter-error progress depends on both nuclear support and update alignment. We then analyze convergence under heterogeneous row-block curvature. Muon's spectral constraint limits the update energy in each row, yielding bounds that depend on average row curvature, whereas the corresponding nSGDM bounds depend on the maximum. We establish stationarity guarantees for exact polar and fixed-iteration Newton–Schulz updates. Under heterogeneous native row-block curvature, the optimized bound is tighter for exact polar than for nSGDM; we also identify conditions under which this advantage extends to the five-step update.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.