acceptodds
Under review as a conference paper at ICLR 2027

Vision Transformers Need Cross Aggregation for Patch Semantics

Abstract

A Vision Transformer (ViT) is tasked to learn both a useful global representation through the CLS token and distinct local semantics through patch tokens. Yet a loss on the CLS token does not always induce learning of discriminative patch tokens. We theoretically observe that a CLS-based loss induces low rank patch gradients and in practice, most patches receive similar gradients. Adding patch-based losses may not compensate for the strong unidirectional CLS induced gradients. Conse- quently, even under global-local objectives, patch updates can remain insufficiently differentiated and spatially correlated. We therefore introduce cross-aggregation, which redefines the CLS token by aggregating patch-level concept heads in the expanded FFN space of every block. This parameter-free operation adds only 0.6% FLOPs to ViT-S and more directly couples global semantics to patch tokens. Combined with local depth-wise convolutions, it further improves the patch seman- tics, in a variety of pre-training methods. Together, these components form the Cross Aggregation and Local Mixing Transformer (CALiT). CALiT consistently improves semantic probes under both global-only and global-local pre-training objectives and vision-language finetuning. In particular, CALiT improves thin class semantics mean IoU by double digits in many pre-training/linear probe pairs and provides moderate gains in all other local semantics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.