acceptodds
Under review as a conference paper at ICLR 2027

Attention Sensitivity in Vision Transformers: Analysis and Robust Small-Data Training

Abstract

Vision Transformers (ViTs) have achieved strong performance in large-scale computer vision tasks but often generalize poorly when trained on small datasets. Recent studies suggest that flat attention is associated with poor generalization, yet how attention patterns affect local sensitivity to perturbations remains insufficiently understood. To study this relationship, we analyze how bounded input shifts induce perturbations in attention logits and use this formulation to characterize attention-level sensitivity. Empirically, models trained on small datasets exhibit higher sensitivity to such perturbations and flatter attention patterns than models pre-trained on large-scale datasets. We further provide a local sensitivity analysis showing that the softmax Jacobian vanishes as attention distributions approach a sharp, one-hot pattern, thereby attenuating the first-order loss variation propagated through attention. This analysis also clarifies that temperature-based sharpening alone does not necessarily reduce sensitivity. Building on this local analysis and the empirical sensitivity gap, we propose Gradient-Guided Robust Attention Sharpening (GRASP), a training-time regularization framework that enforces prediction consistency under gradient-guided attention perturbations and implicitly encourages sharper attention patterns. Extensive experiments demonstrate that GRASP consistently improves ViT generalization in small-data training and out-of-distribution evaluation, while providing additional gains in target-domain fine-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.