What Artifact Tokens in Vision Transformers Are, and What They Do
Abstract
Vision Transformers develop a handful of high-norm “artifact” tokens that occupy low-information image regions and absorb a disproportionate share of attention. Register tokens were introduced to eliminate them, and a subsequent audit reported that the phenomenon does not generalise uniformly across architectures. We present a measurement study across 15 pretrained checkpoints and 10 datasets spanning 9 visual domains. First, artifact formation is determined by model weights, not input statistics, as varying the fraction of low-information patches over a range changes artifact strength with a coefficient of variation of only –, while leaving the peak-artifact layer exactly invariant for every model on every dataset, including histopathology, chest radiography, and satellite imagery. Second, the behaviour emerges with scale, appearing between 22M and 87M parameters on matched DINOv2 checkpoints. Third, register tokens work by capturing the read target. Adding four registers collapses attention mass on high-norm patch tokens from to , renders those patches causally inert under ablation, and makes the registers load-bearing instead. Fourth, for image–text contrastive models, removing eight of roughly 200 tokens costs 15–21 accuracy points under every pooling operator tested, with matched random-ablation controls showing no effect. Fifth, the sign of the ablation effect depends on the training objective rather than the evaluation protocol. It remains invariant across a band of layers, three evaluators, and neighbourhood sizes spanning two orders of magnitude, while exceeding matched random controls by –. This resolves an apparent protocol dependence in which pooled and per-token evaluation yield opposite-signed conclusions about the same intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.