Mirror Descent at the Edge of Stability
Abstract
In many deep learning settings, gradient descent has been shown to operate at the edge of stability – a regime where the largest eigenvalue of the loss Hessian hovers around . We study this phenomenon in mirror descent, a generalization of gradient descent. As mirror descent is non-Euclidean, existing explanations of the edge of stability phase do not extend to this setting. We generalize the existing notion of sharpness to introduce mirror sharpness, which measures the curvature of the loss with respect to the mirror potential. Empirically, we observe that the mirror sharpness hovers slightly above the threshold. Furthermore, we derive a dual central flow ODE for mirror descent that models iterates in continuous time at the onset of edge of stability. Our ODE provides insights on the difficulties of extending existing arguments for why sharpness stabilizes at from gradient descent to mirror descent in full generality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.