Momentum for Variational Learning at Scale
Abstract
Momentum is well-studied for deep learning, but the same cannot be said for variational learning at scale. Here, we revisit the current practices of momentum for variational learning and propose a new momentum where an exponential moving-average of the site functions is maintained. We show that this site-momentum not only generalizes several existing (non-Bayesian) momentum approaches for deep learning (such as MARS and SGDHess) but also speeds up variational learning at scale. For instance, for training Llama 3 (1.3B) on 100B tokens from FineWeb, site-momentum yields a lower perplexity of up to 0.76 throughout compared to the state-of-the-art method. Overall, our work provides new theoretical connections as well as practical improvements for variational learning with momentum.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.