Exits Need Only Classify: Efficient Early Exiting for Vision Transformers
Abstract
Anytime Early Exit (EE) Vision Transformers (ViT) adapt inference to image difficulty along an accuracy-speedup Pareto frontier. Transformer-based exits are the most accurate option but incur excessive per-exit computation and parameter overheads, so the field has turned to cheaper, weaker exits. We show this overhead is avoidable, based on one key insight: unlike an encoder layer that refines tokens for later layers, an EE classifier has one job—to pool and summarize global class information. Three optimizations follow directly. First, because only the [CLS]-token is relevant for image classification, an exit can compute the [CLS]-token update exactly at a fraction of a full encoder update's cost. Second, image patches that are less relevant to classification can be pruned, trading small accuracy losses for additional speedup. Third, because all exits have this one objective, carefully constructed weight sharing allows for reduced exit parameters with small accuracy loss. We instantiate these in Class-only ANytime image Transformers (CANiT) on a frozen ViT backbone. On ImageNet-1k, CANiT improves accuracy over prior EE work by 0.88–3.35 percentage points at matched speedup, with similar gains across three additional datasets. Ablations isolate the contributions of [CLS]-token-only computation, patch pruning, and weight sharing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.