Learning to Allocate Tokens for Vision Transformers
Abstract
Vision Transformers (ViTs) are a critical backbone of modern computer vision frameworks. Standard ViT architectures work by representing uniformly sized patches as single tokens in a sequence. However, this design is poorly aligned with the true nature of images, where information density is typically not spatially uniform. This results in uninformative image patches receiving the same token allocation as highly informative image patches. To address this limitation, we propose the APE (Adaptive Patch Encoder) architecture, an extension of the traditional ViT that estimates patch importances and dynamically allocates and merges tokens to maximize the information carried in the ViT input sequence. We show that APE is highly efficient while achieving state-of-the-art performance on benchmarks across classification, detection, segmentation, and VQA. Quantitative results across a diverse range of tasks validate that APE's learned saliency metric maximizes the compute-performance tradeoff without domain-specific degradation. Checkpoints and code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.