One Score Term Too Many: The Apparent Cost of Normalizing Frozen ViT Tokens Is Largely an Anomaly-Score Artifact
Abstract
Detectors that reconstruct frozen vision-transformer tokens usually L2-normalize each patch token first, and on periodic manufacturing texture this convention appears to be expensive: in a strictly paired counterfactual (one recipe, one boolean flipped) normalization costs +15.38, +38.64 and +23.04 mean anomaly-detection points (mAD) on wafer, SEM-circuit and PCB benchmarks. We show that this effect is largely a property of the anomaly score, not of the features or the objective. The family's score adds a scale-sensitive L2 residual to a cosine one, so normalizing or rescaling the reconstruction target reshapes the anomaly map. Rescoring the same checkpoints with the cosine residual alone, without retraining, removes the effect on all six datasets: on wafer the normalized arm moves from 49.39 to 66.77 mAD, level with the un-normalized arm (65.87), and a fixed-constant rescale that appeared to cost -12.16 mAD costs -0.04. A squared-error loss term also carries scale, by 2-4 mAD at every weight tried, with or without a cosine term beside it. Two public cosine-scored detectors are immune to the identical rescale; one of them, scored with an L2 residual instead, loses 25 mAD over three seeds, and the published UniAD checkpoint, scored with L2, collapses under a 50x test-time rescale that a cosine score survives. The rule is one line: score, and compare, reconstruction detectors under a scale-invariant residual; a one-run diagnostic exposes the artifact.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.