Reconstruction, Not Localization: A Mechanistic Account of Sparse-Feature OOD Detection in Vision Transformers
Abstract
Sparse autoencoder (SAE) features have been proposed as an interpretable basis for out-of-distribution (OOD) detection, inviting the reading that SAEs contain localized "OOD units". We test the controls such a claim needs on frozen vision transformers (13 OOD sets, CIFAR-10-C, a 100-class in-distribution (ID) dataset, and ImageNet-1K scale). In both ReLU and exact-k TopK SAEs (three seeds each), frequently firing units are seed-idiosyncratic, fixed sets of up to five units add about one point of confidence collapse on ImageNet-O beyond the plain reconstruction round-trip, and the largest-contribution units are in-distribution anchors. The collapse is instead the SAE encode->decode round-trip: it lowers a frozen ImageNet head's confidence on ImageNet-O by 70% versus 11% on ID (47% with its native CLS input), whereas matched noise moves it by only 7%; a PCA round-trip reproduces the collapse (73%). It is head-dependent: at a fixed ID, the drop grows monotonically with the number of head classes (from -3% at 10 to 70% at 1,000), while heads trained on the ID label space stay inert (4.5%). We call this Mismatched Wide-Head Reconstruction Sensitivity; it persists with heads trained at each width and on a second backbone (DeiT-B/16). Against published baselines, SAE-based detection is competitive but not better, and a PCA basis matches it. We release a six-step Sparse Reconstruction Audit with margin-reported decision criteria, all JSON artifacts, and a runnable demo; on a released third-party SAE trained on ImageNet it finds no such effect, and a controlled comparison at the same hook ties the effect to an SAE fit on the audited ID. The contribution is not a stronger detector but a tested boundary on what sparse-feature OOD analysis can support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.