RepAlign: A New Regularizer for Adversarial Training
Abstract
Deep neural networks are known to be highly vulnerable to adversarial examples, making robust training a practical necessity. However, robust training often leads to severe accuracy-robustness trade-offs. Existing methods like TRADES and Adversarial Logit Pairing (ALP) align softmax or logit outputs of adversarial inputs with those of clean samples, but are limited to layers at the end of the network. We propose RepAlign, a framework which uses representation similarity (RS) metrics such as Centered Kernel Alignment (CKA) to optimize for alignment between arbitrary features of clean and adversarially perturbed data. The use of general RS metrics allows RepAlign to align features having different properties. Evaluations across CIFAR-10, CIFAR-100, and Imagenette demonstrate that RepAlign consistently pushes the accuracy-robustness Pareto frontier forward. Soft-NN on CIFAR-10 achieves 86.90% clean accuracy (a gain over TRADES) while maintaining 49.37% robust accuracy ( over PGD-AT). On CIFAR-100, Soft-NN strictly outperforms TRADES on both fronts, achieving 60.42% clean and 26.10% robust accuracy (gains of and over TRADES, respectively). On Imagenette, CKNNA closes the gap to clean training () more than any other defense, reaching 88.46% clean accuracy ( over PGD-AT and over TRADES) while sustaining 61.89% robust accuracy. CKA heatmaps show that RepAlign mitigates representation drift and restores the layer specializaion disrupted by standard adversarial training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.