When Does Noise Sensitivity Reveal Head-Accessible Shortcuts?
Abstract
Can mild input noise reveal which coordinates a frozen vision model uses as shortcuts? We measure the class-conditional change of each pooled representation coordinate and use this score in Noise-based Class-Conditional Sensitivity Pruning (NCSP). The method masks sensitive connections in each row of a linear classifier and refits the unmasked head parameters on clean labeled data. The score directly measures brittleness; interpreting it as a shortcut signal additionally requires nuisance-linked coordinates to be head-accessible, coordinate-aligned, and more probe-sensitive than stable semantic alternatives. Under sub-exponential tails, empirical top- masks nearly minimize retained sensitivity mass; the constrained refit is convex and norm-controlled; and cross-row conditions yield prediction and KL-stability bounds. The full pipeline raises BAR accuracy from to , Waterbirds WGA from to , MetaShift WGA from to , and FB-CMNIST accuracy at from to . Matched controls qualify these gains: clean-data head refitting accounts for a large share of the improvement, sensitivity masks do not uniformly beat random or magnitude-based masks, and a 20-class stress test shows no sensitivity-specific advantage. NCSP is therefore a conditional audit and intervention for a particular representation and probe, not a general guarantee of worst-group performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.