Normalization Rigidity and Recovery in Matrix Momentum
Abstract
Matrix optimizers such as Muon normalize a momentum matrix and then change its singular values before updating the model. We study how this transformation responds to gradient noise. Symmetric noise can cancel when the true gradient is zero, while the average update still responds in the wrong direction when a gradient is added. We show that preventing this behavior for every symmetric noise distribution is equivalent to monotonicity of the update map. For methods that apply the same smooth function independently to Frobenius-normalized singular values, this requirement forces the function to be linear once there are at least three singular values. We explain how the shared normalization causes the problem and which gradient directions are affected. We also construct smooth convex tasks on which actual momentum updates keep the gradient away from zero over a rigorously verified time window. Convex spectral potentials provide another design choice: they preserve alignment with the gradient while still bringing singular values closer together. A smoothed nuclear norm additionally controls sensitivity to numerical errors. In our transformer experiments, however, all task-response estimates remain positive, and matched training favors Schatten over the smoothed map. These results identify a mathematical failure mechanism and its limits as an explanation of training behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.