Adaptive Optimization via Momentum on Variance-Normalized Gradients
Abstract
We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization. MVN-Grad scales each coordinate by an exponential moving average of gradient uncertainty and applies momentum to the resulting normalized gradients, removing the cross-time coupling between stale momentum and a stochastic normalizer present in standard Adam-type updates. We prove that, when the mean estimate tracks the conditional mean exactly and the noise is symmetric, this decoupling yields no larger one-step conditional update variance than momentum-then-normalize variance methods, and that, as in LaProp, an isolated spike is normalized before it enters the momentum buffer. In low-variance regimes, we further show that variance normalization avoids sign-type collapse of second-moment scaling in idealized limiting dynamics. Beyond these comparisons, we prove a general nonconvex convergence guarantee for MVN-Grad under bounded-gradient stochastic assumptions. On CIFAR-100 and GPT-style language modeling with models up to 350M parameters, MVN-Grad matches or improves on Adam, AdaBelief, and LaProp, with improved robustness in several evaluated hyperparameter sweeps at the cost of one additional optimizer-state tensor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.