acceptodds
Under review as a conference paper at ICLR 2027

Convergence Analysis of Normalized Gradient Descent for Definable Functions

Abstract

Normalized gradient descent (NGD) updates model parameters in the negative gradient direction with a prescribed step length. Its invariance to positive rescaling and its ability to move through regions with small gradients make it an attractive alternative to gradient descent. We study when diminishing stepsizes make NGD iterates bounded and convergent. For locally Lipschitz definable objectives and every positive vanishing nonsummable schedule, a bounded NGD trajectory has convergent objective values and only Clarke critical accumulation points. For a coercive definable loss, all trajectories of NGD from a bounded initial set are bounded when all the stepsizes are below a uniform threshold. We also prove that stepsizes , with , guarantee convergence of every bounded trajectory to a Clarke critical point for locally Lipschitz losses definable in a polynomially bounded o-minimal structure. This class includes squared-loss objectives for finite ReLU networks on finite datasets. We also show a sharp dependence on the stepsize schedule and dimension: for every and , a real-analytic semialgebraic objective in three dimensions admits a bounded nonconvergent NGD trajectory with . In two dimensions, every bounded NGD trajectory converges throughout for locally Lipschitz definable losses. Thus, even when the loss converges to a critical value and the distance to the set of critical points tends to zero, parameter convergence can depend on dimension and stepsize decay.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.