acceptodds
Under review as a conference paper at ICLR 2027

Attack-Free Adversarial Vulnerability Assessment via Second-Order Statistics of Predictive Probabilities

Abstract

Deep neural networks are widely deployed in safety-critical applications such as medical image analysis and autonomous driving, where adversarial robustness is essential. Existing robustness evaluations primarily measure overall performance degradation under attack, offering limited insight into class-wise or directional vulnerability patterns. We propose a second-order statistical framework that treats class-wise softmax outputs as random variables and analyzes their conditional variance–covariance structure to characterize structural vulnerability. Using only clean validation data-without generating adversarial examples or relying on specific attack algorithms-the framework enables attack-agnostic evaluation while significantly reducing computational cost, serving as a practical pre-screening tool for rapid model assessment. The conditional variance of the correct-class probability quantifies relative class-wise adversarial vulnerability, while conditional covariance reveals directional misclassification tendencies. Experiments on standard benchmarks under diverse untargeted and targeted attack settings demonstrate strong alignment between the proposed metrics and class-wise attack success rates, suggesting that adversarial vulnerability can be effectively characterized without explicit attack generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.