acceptodds
Under review as a conference paper at ICLR 2027

Which features remain vulnerable in robust models?

Abstract

Scaling adversarial training to larger models and hundreds of millions of synthetic images has improved robustness, yet these classifiers remain vulnerable to small input perturbations. What properties of their learned features make these attacks possible? We study 71 of the most robust CIFAR10 models publicly available, spanning a wide range of architectures and training conditions. Our analysis reveals that adversarial attacks manipulate the same features that are responsible for separating the classes on clean data. Moreover, this relationship is concentrated in a surprisingly small subset of features: for the most robust model, more than 70% of successful attacks can still succeed when restricted to just 10% of the features per class pair. Furthermore, we find that the samples that remain vulnerable can be identified as those which are poorly separated along these specific feature dimensions; and that successful attacks follow a consistent attack direction. In addition, at the model level, we demonstrate a gradient-free measure of margin spread that is strongly correlated with higher adversarial robustness. Together, these results provide new insight into the mechanisms that underlie the remaining vulnerability of highly robust classifiers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.