acceptodds
Under review as a conference paper at ICLR 2027

Last-Iterate Convergence Rate of Normalized Gradient Descent under Hölder Smoothness

Abstract

Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study the last-iterate convergence of normalized gradient descent for convex, -Holder-smooth objectives. For a constant stepsize, we establish an upper bound of , which contains a logarithmic overhead relative to the known guarantees for the best and weighted-average iterates. For , this overhead is known to be unavoidable. We complement this analysis with numerical results based on the performance estimation problem (PEP), investigating the finite-horizon worst-case behavior in the smooth setting and whether the logarithmic overhead reflects an intrinsic limitation of constant-step normalized gradient descent. We then show that a linearly decreasing stepsize yields a last-iterate guarantee of , matching the order of the best-iterate/weighted-average guarantees without requiring knowledge of and .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.