Beyond Singular Values: Geometry and Numerical Fidelity in Muon
Abstract
When does spectral flattening improve the loss decrease from a matrix update? We study this question for Muon through rotations that preserve the candidate spectrum, update norm, and entire ideal polar update while changing the orientation of a spectrum-preserving comparator. On one fixed objective, candidates with identical spectra and gradient alignments can favor opposite updates; a local theorem identifies the curvature projection responsible for this variation. At finite step sizes, the gradient at the comparator endpoint determines the first-order response to a small rotation and accounts for curvature along the update path. Tests span 30 frozen state–block combinations from three 124-million-parameter language-model trajectories. In the first layer, scalars fitted on old text recover most of the endpoint predictor's gain. In middle and final layers, the same-objective endpoint expansion remains accurate, while calibration-to-evaluation differences dominate prediction error. We then separate the exact five-round spectral map from its floating-point implementation. In a first-layer same-input test with raw output lengths, single precision (FP32) reduces paired-response squared error by 99.97% compared with brain floating point (BF16). FP32 preserves the ideal response sign in all 36 first-layer and 24 deeper-layer constructed directions; the high-precision finite-round map reproduces the remaining random-direction reversals. These results distinguish orientation and path effects from data-transfer error and numerical distortion at the intervention scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.