LOVA: Combating Local Variance Dominance for Efficient Neural Network Training
Abstract
Adam is one of the most dominant optimizers for training neural networks. Nevertheless, it has been challenged by matrix-based optimizers recently. In this work, we identify that the signal is often dominated by the random noise in Adam, an issue that becomes increasingly severe in large-batch training. The consequence is catastrophic—Adam overdampens the updates and heavily impedes the training efficiency. To address this issue, we propose Lova (Local Variance Adaptation), a memory-efficient optimizer that preserves the signal drift while retaining the adaptivity of coordinate-wise learning rate. It should be highlighted that Lova tracks only the first-moment buffer, and the memory footprint is as low as stochastic gradient descent (SGD) with momentum. While Adam is not guaranteed to converge for smooth nonconvex functions, theoretical analysis suggests that Lova, under standard assumptions, achieves a convergence rate of . Importantly, this rate holds for an arbitrary constant and eliminates the polynomial dependence on the reciprocal of a small constant for numerical stability, . Extensive experiments on various architectures, datasets, and tasks demonstrate the superiority of Lova over AdamW and other popular optimizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.