Does Certified Robustness Generalize? A Neural Network Verification Study of Distribution Shift
Abstract
Certified robust accuracy (CRA) is the standard metric for quantifying worst-case adversarial robustness and for benchmarking certified defenses. However, CRA is by construction reported only on the in-distribution (ID) test set on which the model was trained. A largely understudied question is what CRA tells us about broader model behavior, in particular predictive accuracy, and whether the relationship between certified robustness and accuracy carries a transferable structure under distribution shift. This is non-trivial because formal verification is strictly stronger than any single-attack evaluation: a CRA guarantee asserts the absence of any adversarial example in an ℓ∞-ball, so the “on-the-line” transfer observed for clean accuracy and empirical attacks need not carry over. We present, to the best of our knowledge, the first systematic study of CRA under distribution shift for verification-tractable architectures, covering MNIST-FC, GTSRB-CNN, and CIFAR-CNN across structural and corruption shifts. Using clean-accuracy-matched model families and shared-property verification, we find that (i) the inter-model CRA ordering shows strong but imperfect agreement from ID to OOD, with statistically significant ID advantages retained under the tested shifts, and (ii) ID CRA correlates with OOD accuracy, calibration, and OOD detection across a multi-architecture ablation, acting as a model-selection proxy for several practical OOD properties. We further support these findings with a machine-checked theoretical framework that characterizes when ID CRA bounds OOD performance and with controlled-shift experiments that test the bound directly. Selecting by ID CRA incurs 0–7 percentage points of OOD CRA loss in the tested settings. These findings extend CRA’s value beyond robustness benchmarking, supporting its use as a model-selection signal for accuracy, calibration, and detection performance under distribution shift. Our code is available at [https://anonymous.4open.science/r/veri-generalization-anonymous](https://anonymous.4open.science/r/veri-generalization-anonymous).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.