acceptodds
Under review as a conference paper at ICLR 2027

SALF: Spatially Attentive Layer Fusion for Probing Vision Transformers

Abstract

Adapting a pre-trained Vision Transformer (ViT) to a downstream task requires extracting a single compact representation from multi-layer patch tokens. Current methods either attentively probe the patches of the final layer (discarding intermediate features) or fuse per-layer average-pooled summaries across depth (discarding fine-grained spatial detail). By probing all patches across all layers on multiple datasets, we demonstrate that the final layer is rarely the most informative and that dynamic spatial aggregation significantly outperforms static average pooling. Notably, our analysis reveals that this spatial aggregation plays a more dominant role in downstream performance than depth-wise fusion, effectively reducing the reliance on complex multi-layer integration. To efficiently leverage both dimensions, we factorize all-token readout into SALF-P (Spatially Attentive Layer Fusion with Per-layer aggregation), which first aggregates the spatial patches of each layer using a parameter-efficient shared cross-attention module, followed by a second cross-attention module to fuse these features across layers. However, further analysis reveals that depth-wise fusion has limited impact when paired with spatial attention. This motivates our SALF-A, which averages over layer features before applying a single attentive pooling. SALF-A matches SALF-P across our benchmarks, while reducing the spatial readout to one pass and shrinking the token cache by a factor of the network depth. Evaluations across 19 downstream datasets and seven frozen backbones, alongside Image Quality Assessment (IQA) tasks under parameter-efficient fine-tuning, confirm that SALF-P and SALF-A substantially improve accuracy on fine-grained visual perception tasks by up to 16.20 percentage points over ALF while maintaining strong performance on globally semantic tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.