FAST ADVERSARIAL TRAINING NEEDS DEBIASING AND CORRECTION TO ESCAPE THE TRAP
Abstract
Fast adversarial training (FAT) reduces the cost of robust training through single- step perturbation generation but remains vulnerable to catastrophic overfitting (CO), an abrupt collapse in adversarial robustness. By systematically dissecting CO from two complementary perspectives, we uncover two key failure mechanisms. First, we uncover an implicit class-specific backdoor-like effect: the outer minimization repeatedly associates biased perturbations with their ground-truth labels, causing persistent class-specific patterns to evolve into predictive triggers. Second, we un- cover gradient drift along the often-overlooked post-initialization update trajectory, where the initialization gradient becomes severely misaligned with subsequent local loss-ascent directions. To address these failures, we propose DC-FGSM, which combines class-adaptive spatial debiasing to disrupt the accumulation of class-specific trigger patterns with drift-aware gradient correction to compensate for directional deviations. Experiments across multiple datasets and perturbation budgets demonstrate that DC-FGSM achieves state-of-the-art robustness among single-step adversarial training methods. Furthermore, our gradient correction mod- ule consistently improves a broad range of existing FAT baselines, demonstrating its plug-and-play generality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.