ATFeaturePure: Learning A Class-Conditional Vector Field in Decision-Relevant Feature Space for Adversarial Defense
Abstract
Adversarial training (AT) provides a reliable foundation for adversarial robustness, yet adversarial perturbations can still substantially corrupt the internal representations that determine model predictions. Rather than recovering the entire clean input with a high-dimensional generative purifier, we investigate whether AT can be strengthened by directly correcting these decision-relevant representations. We introduce , which retains AT as the robust learning foundation and augments the classifier with a lightweight class-conditional vector field operating on the compact pre-logit feature. The key challenge is how to learn correction dynamics that actually improve classification robustness. To address this, we develop a local theoretical analysis that identifies the behavior required for effective feature correction and translate these insights into a learning mechanism that jointly trains the classifier and vector field while encouraging controlled local sensitivity, reliable class guidance, correction of adversarial features toward clean representations, and decision-consistent predictions. Extensive experiments across three datasets, two backbones, and multiple strong attacks show that ATFeaturePure consistently improves robustness over recent AT and purification baselines, including against adaptive attacks, while maintaining favorable training efficiency. These results show that theory-guided correction of decision-relevant features can effectively strengthen AT without requiring full input reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.