Augmenting Matrix Orthogonalization with Sample-Space Geometry
Abstract
Orthogonalized optimizers such as Muon exploit the matrix structure of parameter updates to improve neural network training. Their normalization acts on the row and column directions of aggregated updates, leaving the co-update directions shared across individual training examples without explicit spectral control. We propose augmenting matrix-based orthogonalization with this complementary sample-space geometry. By organizing per-sample updates into a matrix, we identify shared update directions through a sample-space Gram matrix, enabling spectral shaping across vector, matrix, and heterogeneous parameter blocks. To incorporate these directions without disproportionately amplifying weak components, we introduce Selective Orthogonalization: a smooth transformation that normalizes strong singular directions while preserving the scale of vanishingly weak ones. We prove that, for a fixed positive transition scale, its relative amplification approaches one as a singular value vanishes, and derive an equivalent parameter-side spectral formulation. The resulting approach supports a hybrid optimizer that alternates Selective and Muon updates across iterations, bringing sample-space and parameter-matrix geometry into the same training process. Standalone Selective experiments establish the behavior of the sample-space component: controlled studies validate weak-direction scale preservation across five parameter structures, while image-classification experiments demonstrate substantially reduced weak-direction amplification and competitive accuracy across datasets, architectures, and batch sizes. Together, these results motivate per-sample update geometry as a complementary source of structure for orthogonalized optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.