acceptodds
Under review as a conference paper at ICLR 2027

Momentum Acceleration of Normalized Steepest Descent at the Edge of Stability

Abstract

Optimizers based on normalized steepest descent with momentum are widely used to train large language models, yet existing analyses attribute the benefit of momentum to stochastic noise reduction and favor no momentum in deterministic training. We show that full-batch training nevertheless benefits from momentum and explain why. Once training reaches the Edge of Stability, normalization pins the effective learning rate at twice the inverse curvature transverse to the gradient, so progress is limited by the loss geometry rather than by the learning rate. Momentum attenuates the alternating gradient component before normalization, which raises the threshold at which oscillations start and multiplies the progress after it by . A switched flow built from this mechanism predicts the onset of oscillations, the speed afterward, and, integrated from initialization, the training loss in full-batch experiments. We prove the mechanism on high-dimensional quadratic valleys, including a strict speedup from the same initialization without assuming periodic behavior, and show that on general smooth losses normalized descent follows the flow locally once it oscillates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.