ACTIVATION-MASK STABILITY AND ADVERSARIAL VULNERABILITY: A MEASUREMENT-COUPLING CASE STUDY
Abstract
A statistic measured at an adversarial example can appear extraordinarily predictive of that example’s adversarial distance even when it carries little independent information, simply because the measurement point is chosen by the same search that defines the distance. We demonstrate this for activation-mask stability DM(x, δ) = ∥M(x + δ) − M(x)∥0/N in ReLU networks. Across five sparsity-regularization strengths and three datasets (MNIST, FashionMNIST, CIFAR-10, plus a PGD-AT CIFAR-10 control), per-example activation sparsity S(x) is a weak-to-null correlate of adversarial distance R(x) (r ∈ [−0.20, 0.31]), whereas DM measured at the minimal perturbation found by DeepFool is strong (r = 0.91–0.98). But R(x) = ∥δ∥2 is exactly the distance to that same point, so the two are coupled by construction. Using the same 20 checkpoints and no new experiments beyond what we already ran, we test this three ways: two δ-free alternatives— fixed-radiusmask sensitivityDϵM (x) and clean-input boundary distance AFC(x) — both collapse relative to the attack-point measurement (r = 0.10– 0.17 and near-zero to −0.23, respectively); and a crossing-density decomposition C(x) = DM/R testing whether DM is merely proportional to distance traveled, which holds on MNIST (r = −0.04) but not the other three datasets (r = −0.11 to −0.35). Together these support reading most, but not all, of the attack-point correlation as measurement coupling rather than an independent relationship of comparable size, with a smaller residual (|r| ≈ 0.1–0.35) surviving every way we decoupled it and reproducing under an independent CW-L2-style attack. We frame this as a case study of one measurement-coupling mechanism in one small ReLU architecture, not a general claim about activation sparsity or mask-stability metrics at large.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.