acceptodds
Under review as a conference paper at ICLR 2027

Exits Need Only Classify: Efficient Early Exiting for Vision Transformers

Abstract

Anytime Early Exit (EE) Vision Transformers (ViT) adapt inference to image difficulty along an accuracy-speedup Pareto frontier. Transformer-based exits are the most accurate option but incur excessive per-exit computation and parameter overheads, so the field has turned to cheaper, weaker exits. We show this overhead is avoidable, based on one key insight: unlike an encoder layer that refines tokens for later layers, an EE classifier has one job—to pool and summarize global class information. Three optimizations follow directly. First, because only the [CLS]-token is relevant for image classification, an exit can compute the [CLS]-token update exactly at a fraction of a full encoder update's cost. Second, image patches that are less relevant to classification can be pruned, trading small accuracy losses for additional speedup. Third, because all exits have this one objective, carefully constructed weight sharing allows for reduced exit parameters with small accuracy loss. We instantiate these in Class-only ANytime image Transformers (CANiT) on a frozen ViT backbone. On ImageNet-1k, CANiT improves accuracy over prior EE work by 0.88–3.35 percentage points at matched speedup, with similar gains across three additional datasets. Ablations isolate the contributions of [CLS]-token-only computation, patch pruning, and weight sharing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.