Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
Abstract
We establish the first convergence guarantees of gradient descent for general feedforward neural networks of any width or depth, with any initialization or dataset. We only assume that the activation functions are linearly bounded, Lipschitz continuous, and Lipschitz smooth—properties that hold for linear, tanh, softplus, sigmoid, and smoothed ReLU functions—and that the loss function is Lipschitz smooth in the model outputs, a mild condition satisfied by mean squared error and binary cross-entropy loss. By relating the Lipschitz properties of one layer to the next, we obtain a novel generalized Lipschitz smoothness condition for an -layer neural network where the change in gradient is upper bounded by the change in the parameter space, multiplied by a polynomial of the parameter norms at both endpoints of degree . This yields a descent lemma where the loss decreases as long as the learning rate is small enough with respect to the parameter norms. By ensuring that the leading term of the polynomial grows at a controlled rate, we prove that the minimum squared gradient norm converges to zero in iterations at rate .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.