Reference Saturation: Normal Exemplars Buy Structural Anomaly Detection Twice as Fast as Logical
Abstract
Zero- and few-shot anomaly detectors are almost always reported on MVTec LOCO as a single pooled AUROC, even though the benchmark separates structural anomalies, whose local appearance is atypical, from logical anomalies, whose every local region looks normal and only the composition is wrong. We re-measure the reference-based family under one protocol, split by subset, and find a stable asymmetry: going from zero to thirty-two normal exemplars closes 41.9% of the remaining headroom on structural anomalies but only 20.7% on logical ones, a ratio of 2.02 (95% CI [1.67, 2.51]), which an independent redraw of the thirty-two exemplars reproduces at 2.02 [1.69, 2.51]; it is 2.00 at sixteen exemplars and, for a second method with a different mechanism measured from its one-exemplar baseline, 1.75, and every interval excludes one. Reference diversity then saturates, and reference depth recovers little of what it leaves. Adding a third reference branch to a two-branch combination moves logical AUROC by at most 0.2 points while roughly doubling the exemplar budget, and spending the entire budget on one method's exemplar diversity buys nothing measurable: three independently seeded WinCLIP+ branches at 46.4 exemplars reach 66.4, against 66.1 for a single branch at 16 and 69.6 for the two-branch combination at that same 16; pooling the same budget into one 32-exemplar memory per branch instead, the strongest reference-only configuration we could build, reaches 70.5, a gain of 0.7 over the best diverse one that we cannot resolve. Replacing that third branch with a compositional part-count score — whose per-image scores on logical anomalies agree with the reference branches far more weakly than they agree with each other — raises logical AUROC from 69.6 to 73.3±1.2 over three independent draws of the composition branch's own references, at no more than 32 images per category. Against that deeper control the paired bootstrap over images gives +2.85, 95% CI [0.97, 4.79], and against the best three-branch reference configuration +3.58, 95% CI [1.77, 5.40]; structural AUROC does not move resolvably against either. With all three branches reading the same sixteen images, half the control's budget, the gain over it is +3.74, 95% CI [1.88, 5.66]. We report a pre-specified control that falsified our own initial two-branch claim, four protocol anchors against published numbers, and a spread of up to 17.8 points across three published tables for the same baseline at the same shot count. All per-image scores, exemplar manifests and per-category budgets are released. We make no state-of-the-art claim.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.