WHEN ATTENTION GEOMETRY SHIFTS: CALIBRATED STRUCTURAL EVALUATION OF VISION TRANSFORMERS
Abstract
Vision transformers route information through attention, yet shift detection often relies only on pooled features or classifier confidence. This paper studies whether a query’s routing pattern can be compared with nominal executions to characterize visual shift. Attention-Geometry Tracking (ATH) forms a calibration bank of clean embeddings and attention summaries, retrieves a coverage-qualified Local reference or falls back to an Aggregate reference, and scores shifts in key marginal, stable rank, entropy dispersion, and mean entropy. It requires neither labels nor classifier logits at scoring time and records the reference branch for every query. On a 75-condition window-averaged Swin-B corruption grid, mean-entropy shift (MES) is the strongest tested standalone score (mean AUROC 0.8049); the evaluated transferred-weight fusion reaches 0.6407. In a four-model OOD experiment—ViT-B/16, Swin-B, DINOv2-B/14, and CLIP ViT-B/16—Local references improve ImageNet-R AUROC over Aggregate references for every model, whereas DTD effects are mixed. Reference conditioning can therefore alter rankings but is shift dependent; high ImageNet-R FPR@95 precludes an operating- point claim. The corruption grids are explicitly exploratory because their selection slices overlap evaluation sources.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.