acceptodds
Under review as a conference paper at ICLR 2027

SAGA: Spatially-Aware Gated Attention for Separating Outlier Structure from Functional Effect in Vision Transformers

Abstract

Vision Transformers can develop a small number of unusually high-norm patch tokens, motivating methods that suppress, relocate or regularize this activity. Yet reducing an outlier count or producing a cleaner spatial map does not necessarily imply that the representation used by a downstream task has improved. We introduce SAGA (Spatially-Aware Gated Attention), a lightweight receiver-side mechanism that modulates patch-token attention outputs by layer, head and spatial position without adding extra tokens. Across three audited 300-epoch ImageNet-1K architecture-recipe cells with matched Baseline-SAGA checkpoints, SAGA reduces fixed-threshold high-norm-token counts by approximately – while yielding - percentage-point higher observed mean top-1 accuracy than the ViT baselines. Mechanistically, high-norm positions receive - more incoming attention per token than other patch positions and form highly reproducible spatial patterns. Matched frozen interventions, however, show that this spatial reproducibility does not translate into a reproducible coordinate-specific classifier response in the tested blocks. Functional relevance also depends on the readout: terminal patch-gate edits can alter patch representations while leaving CLS logits exactly unchanged, whereas, in a descriptive patch-level analysis, high-norm positions are associated with substantially poorer exact-transform correspondence. These results separate fixed-threshold outlier prevalence, spatial organization and readout-specific functional effect, motivating evaluation of ViT diagnostics against the downstream computation that consumes them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.