acceptodds
Under review as a conference paper at ICLR 2027

Fine-grained Analysis of Adam's Implicit Bias Near the Manifold

Abstract

Understanding the implicit biases of different optimizers is one of the key challenges in studying the training dynamics of deep learning. However, the implicit bias of Adam, one of the most widely used adaptive gradient methods (AGMs), is still not fully understood theoretically. li2025adam studied the implicit bias of Adam with a second-moment momentum coefficient , a setting they termed the 2-scheme. They proved that, under the 2-scheme, Adam evolves near the manifold of training loss minimizers and implicitly reduces a form of sharpness different from that associated with SGD over an timescale. However, in practice, is much larger than . This leaves an open theoretical question: what is the implicit bias of Adam in the -regime with ? In this work, we resolve this question. We first extend the convergence results of li2025adam to and provide counterexamples showing that such convergence can fail for . We then characterize Adam's implicit bias for by developing a new framework for analyzing the limiting flows of AGMs, including Adam, over long training horizons. This framework generalizes the ideas in katzenberger1991solutions and also provides insight into a broader class of stochastic processes beyond AGMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.