acceptodds
Under review as a conference paper at ICLR 2027

Learning to Allocate Tokens for Vision Transformers

Abstract

Vision Transformers (ViTs) are a critical backbone of modern computer vision frameworks. Standard ViT architectures work by representing uniformly sized patches as single tokens in a sequence. However, this design is poorly aligned with the true nature of images, where information density is typically not spatially uniform. This results in uninformative image patches receiving the same token allocation as highly informative image patches. To address this limitation, we propose the APE (Adaptive Patch Encoder) architecture, an extension of the traditional ViT that estimates patch importances and dynamically allocates and merges tokens to maximize the information carried in the ViT input sequence. We show that APE is highly efficient while achieving state-of-the-art performance on benchmarks across classification, detection, segmentation, and VQA. Quantitative results across a diverse range of tasks validate that APE's learned saliency metric maximizes the compute-performance tradeoff without domain-specific degradation. Checkpoints and code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.