acceptodds
Under review as a conference paper at ICLR 2027

Attention Shrinkage for Vision Transformers via Posterior Weighting

Abstract

Vision Transformers (ViTs) utilize multi-head self-attention to capture global context, yet their raw attention scores are inherently noisy and densely distributed, often causing attention to bleed into irrelevant background elements and thereby limiting model interpretability. Regularization of attention scores is understudied, while existing methods require architectural modification, use annotated segmentation, or aggressively hard-thresholding. To address these limitations, we propose adaptive posterior weighting methods — specifically Posterior Probability Shrinkage (PPS) and Wiener Shrinkage — which act as robust, data-driven regularizers. By penalizing attention values proportionally to their likelihood of belonging to a null distribution or by attenuating noisy spectral components, our approaches probabilistically refine attention maps. To overcome the computational burden of resampling, we introduce a joint bootstrap distribution that aggregates null statistics across the dataset. While architecture-agnostic, our proposed posterior weighting improves sparsity, shrinkage, and interpretability. _Without optimizing for localization or segmentation_, regularized attention maps enhance spatial coherence and smoothness. Evaluated on DUTS (a subset of the ImageNet) and LIDC-IDRI (thoracic CT scans) using three ViT architectures, our automated calibration and posterior weighting significantly improve the localization of target regions. Overall, the proposed regularization techniques for attention scores offer a superior balance by producing substantially smoother, less noisy attention distributions, without any hyperparameter tuning. These benefits are pronounced in architectures with inherently noisier attention mechanisms, demonstrating a statistically principled and robust path toward more interpretable visual representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.