acceptodds
Under review as a conference paper at ICLR 2027

Continuous Approximation for Gradient-based Optimization: Arbitrary-Order Discretization Error and Implicit Bias

Abstract

Gradient-based optimization methods are the backbone of modern machine learning and deep learning, and thus the theoretical understanding of their dynamics and implicit biases has been crucial recently. A promising line of works in this direction leverages continuous-time approximations to capture the behavior of discrete and iterative updates, shedding light on their convergence, stability, and regularization properties. However, these approximations typically suffer from discretization errors that could become significant with large learning rates. To this end, we propose a unified framework for constructing continuous-time approximations with discretization error to arbitrary order of learning rate. This framework can be applied to a wide range of gradient-based methods, including momentum methods and adaptive variants like Adam. This is achieved by segmented Taylor expansions and parameter-splitting techniques that generalize prior works of standard gradient descent, hence enabling a precise analysis of more delicate optimizer dynamics. As a key case study, we derive the first third-order continuous approximation of Adam, which exhibits novel implicit regularization terms that cannot be obtained from lower-order ones. This allows us to identify a new phase transition occurring in Adam's training, explaining its acceleration and generalization behavior, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.