acceptodds
Under review as a conference paper at ICLR 2027

Normalization Rigidity and Recovery in Matrix Momentum

Abstract

Matrix optimizers such as Muon normalize a momentum matrix and then change its singular values before updating the model. We study how this transformation responds to gradient noise. Symmetric noise can cancel when the true gradient is zero, while the average update still responds in the wrong direction when a gradient is added. We show that preventing this behavior for every symmetric noise distribution is equivalent to monotonicity of the update map. For methods that apply the same smooth function independently to Frobenius-normalized singular values, this requirement forces the function to be linear once there are at least three singular values. We explain how the shared normalization causes the problem and which gradient directions are affected. We also construct smooth convex tasks on which actual momentum updates keep the gradient away from zero over a rigorously verified time window. Convex spectral potentials provide another design choice: they preserve alignment with the gradient while still bringing singular values closer together. A smoothed nuclear norm additionally controls sensitivity to numerical errors. In our transformer experiments, however, all task-response estimates remain positive, and matched training favors Schatten over the smoothed map. These results identify a mathematical failure mechanism and its limits as an explanation of training behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.