Attacking a Linear Probe measures the Head as much as the Encoder
Abstract
Adversarial evaluations of pretrained vision encoders often freeze the encoder, train a linear classifier, and attack the resulting model. How much of the reported robustness belongs to the classifier head? Across seven ImageNet-1k probes, we retrain six heads without their BatchNorm layer while holding the encoder and seed fixed. Under the AutoAttack ensemble, this intervention raises MoCo v3's robust accuracy from 3.88% to 6.26%: a change equal to 30% of the original between-encoder spread. In all six pairs, the sign of the change matches whether the original BatchNorm layer's mean feature-wise scale is below one (attenuation) or above one (amplification), with three differences significant under corrected paired tests. Cross-entropy PGD exaggerates the effect enough to reverse one clear pairwise encoder ordering; the ensemble does not resolve a rank reversal. A second intervention isolates attack sensitivity: dividing a trained head's logits by a positive constant preserves its decision function, yet changes PGD accuracy by up to 28.7 percentage points on MoCo v3. The extreme increase reflects gradient masking, not improved robustness. The two trained heads also differ under an attack with a scale-invariant objective. By contrast, the BatchNorm intervention changes ImageNet-C accuracy by at most 2.0 percentage points, leaving the large gaps between encoders essentially unchanged. These results support controlling the head and checking the attack when using frozen probes to compare representations. They establish dependence on the trained probe, while leaving the roles of head scaling and optimisation only partly separated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.