A Training-Dynamics View of Catastrophic Overfitting: Understanding and Prevention
Abstract
We study the underlying mechanism of catastrophic overfitting, a phenomenon in which models overfit to weak adversarial examples and lose true robustness, from the perspective of training dynamics. While prior work explains *when* catastrophic overfitting occurs through input-space quantities, the training-dynamics mechanism by which parameter updates produce the sudden collapse has received comparatively little attention. Catastrophic overfitting is a failure of *training*, and what corrupts training is the parameter update; we therefore analyze the *mixed Hessian*, the operator that carries an input-space perturbation error into the parameter-update error. We provide analytical and interventional evidence that rapid amplification of the mixed Hessian drives the parameter-update misalignment underlying catastrophic overfitting, and show that directly suppressing its spectral norm prevents the collapse. Based on this insight, we propose a novel KL divergence-based regularizer that stabilizes training dynamics and effectively prevents catastrophic overfitting. Our method achieves the strongest robustness among single-step methods and, when combined with multi-step adversarial training, further improves robustness over multi-step training alone, indicating that mixed-curvature stabilization is a general principle rather than a single-step remedy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.