Gradient Clipping at and Beyond the Edge of Stability
Abstract
When training deep neural networks with full-batch gradient descent, the curvature of the loss stays at the threshold of instability, hovering around where is the learning rate. We study how clipping changes this Edge-of-Stability (EoS) phenomenon in first-order (FO) and zeroth-order (ZO) optimization under batch and per-sample clipping. Our central finding is that clipping stabilizes FO and ZO training in qualitatively different ways. For FO, clipping shifts the relevant stability edge, and the trajectory remains near this clipping-aware edge; stronger clipping permits stability at larger curvature. For batch-clipped ZO, training remains bounded at curvatures well beyond the mean-square stability threshold, even though standard stability analysis predicts instability in this regime. For both batch- and per-sample clipped ZO, we characterize nonzero fixed points through the covariance of the parameter displacement and show local attraction within the corresponding closed covariance dynamics under explicit conditions. Experiments show that FO trajectories track clipping-aware stability edges, while batch-clipped ZO approaches its nonzero covariance balance beyond the predicted stability limit and per-sample ZO moves toward the predicted matrix-valued covariance balance along the leading-curvature directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.