acceptodds
Under review as a conference paper at ICLR 2027

First-Order Characterization of Probabilistic Robustness for Classifiers via the Coarea Formula

Abstract

Adversarial training (AT) is a standard approach for improving worst-case robustness, but its worst-case formulation can be overly conservative when perturbations are predominantly stochastic. This has motivated PR-targeted training, where existing methods typically rely either on surrogate risk objectives or on adaptations of adversarial-search procedures, leaving the first-order behavior of PR itself largely unexplored. In this paper, we study the first-order optimization structure of PR and ask when its variation can be represented by only a few perturbations. Under mild regularity conditions, we express the PR gradient via the co-area formula as a distribution-weighted integral over the decision boundary, revealing an inherently aggregate boundary effect. To study this quantity in high dimensions, we use a Gaussian kernel as a computational probe, yielding a simple Monte Carlo estimator of the PR first variation. Low- and high-dimensional experiments show that the extent to which a small number of perturbations represent the aggregate PR variation depends strongly on the underlying geometry and dataset. We further study how the estimated PR variation depends on the chosen margin level and smoothing bandwidth, and explore its use in PR-oriented training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.