Orthogonalization Separates Sharpness Exploration from Perturbed Gradient Updates
Abstract
In Sharpness-Aware Minimization (SAM), the perturbed gradient updates the model parameters, while the clean gradient determines the parameter perturbation. This raises a natural question: does matrix orthogonalization have different effects when applied to the parameter perturbation versus the subsequent parameter update? We thus investigate how SAM changes when matrix orthogonalization is introduced separately into these two stages: 1) for clean gradient, orthogonalization changes the first order sharpness term from a Frobenius norm quantity to a normalized nuclear norm quantity, providing a nuclear norm view in sharpness exploration; 2) for perturbed gradient, orthogonalization changes how gradient variation caused by the perturbation enters the final update. We further analyze the effects of finite step Newton–Schulz approximation and global normalization on these two mechanisms. With the above analysis, MuonSAM instantiates this construction by orthogonalizing both the clean and perturbed gradients. Experiments on vision and language tasks verify the theoretical predictions and provide evidence for the roles of orthogonalization in sharpness exploration and model updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.