AGAM: Alignment-Gated Momentum for Deep Learning Optimization
Abstract
Momentum-based optimizers underpin modern deep learning training, relying on accumulated past gradients to provide a reliable descent direction. However, this expectation is routinely violated in practice due to the non-stationarity of mini-batch gradients, causing a base optimizer’s proposed update direction to become misaligned with the current gradient. We introduce , a state-preserving, drop-in gating mechanism for momentum-based optimizers, implemented as a modification with minimal computational overhead and no additional hyperparameters. AGAM transiently attenuates the contribution of historical momentum via a tensor-wise, alignment-adaptive gating coefficient. We instantiate AGAM with SGDM, AdamW, and Lion, characterize their exact relationship to quasi-hyperbolic optimization, and establish a sharp descent certificate for AGAM-SGD. Empirical evaluations across ResNet image classification, MAE ViT pre-training, and LLM fine-tuning demonstrate that AGAM-augmented optimizers consistently outperform both their base optimizers and alignment-aware baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.