acceptodds
Under review as a conference paper at ICLR 2027

From the Edge of Stability to Convergence Theory

Abstract

Gradient descent is usually analyzed under smoothness conditions that guarantee a decrease in the objective at every step. In modern neural-network training, however, gradient descent often operates at the Edge of Stability (EoS), where the largest Hessian eigenvalue oscillates around or above the classical stability threshold and the loss is not monotone. We study what controls convergence in this regime. Our analysis uses an averaged directional sharpness that takes into account the Hessian spectrum along the gradient-descent update. We show that a squared-gradient-weighted stability margin built from this quantity exactly determines the cumulative loss decrease along the trajectory. Lower bounds on this margin directly give gradient-stationarity rates, as well as convergence rates under PL condition. To control the margin when the loss oscillates, we introduce a sufficient condition based on comparing iterates steps apart. The condition is derived using classical smoothness, but can still hold when individual gradient-descent steps increase the loss. Importantly, we show that it is satisfied by various theoretical neural-network models. Experimentally, we verify lower margin bounds in MLPs and Vision Transformers. Across full-batch experiments on CIFAR-10 and Tiny ImageNet, the cumulative margins is shown to control the dynamics, and we further characterize how mini-batching affects averaged sharpness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.