Understanding Edge of Stability in Rank-1 Linear Models for Binary Classification
Abstract
Recent research in deep learning optimization reveals that many neural network architectures trained using gradient descent with practical step sizes, , exhibit an interesting phenomenon where the top eigenvalue of the Hessian of the loss function, , increases to and oscillates about the stability threshold, . The two parts of the trajectory are referred to as Progressive Sharpening and Edge of Stability. The oscillation in is accompanied by a non-monotonically decreasing training loss. In this work, we study the Edge of Stability phenomenon in a two-layer rank- linear model for the binary classification task with linearly separable data to minimize logistic loss. By capturing the core training dynamics of our model as a low-dimensional system, we rigorously prove that Edge of Stability behavior is not possible in the simplest one datapoint setting. With two datapoints, empirically we observe that, when the margin between the two points is small, Edge of Stability may occur. The oscillation in loss and sharpness becomes perpetual in the limit where the margin converges to 0. Using bifurcation and normal form analysis of the dynamical system in this asymptotic setting, we provide a sufficient condition under which the dynamical system exhibits perpetual oscillation in both loss and sharpness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.