Anticorrelating Minibatch Noise Accelerates Adam
Abstract
A minibatch gradient can be written as the full-dataset gradient plus a zero-mean sampling residual, which is temporally uncorrelated under independent sampling. This residual is central to stochastic optimization, but its temporal organization across updates remains largely overlooked. We introduce temporal antithesis, which makes consecutive residuals anticorrelated while preserving the minibatch one-step gradient's mean and covariance. We characterize analytically how temporal antithesis alters Adam's dynamics and show empirically that it substantially accelerates Adam per update, especially early in training. Because exact antithesis requires the full-dataset gradient, we propose superbatch constructions that approximate it using several ordinary minibatches. These constructions recover most of the early advantage of exact antithesis. We further observe the same early acceleration across convolutional networks, vision transformers, and a 124M-parameter GPT-2 model, as well as beyond Adam with SGD and Muon. Together, these findings establish temporal noise correlation as a consequential direction for understanding and designing stochastic optimizers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.