acceptodds
Under review as a conference paper at ICLR 2027

Fast Adaptive Token Selection via Efficient Reinforcement Learning

Abstract

Vision Transformers (ViTs) have achieved remarkable success in a wide range of visual tasks due to their strong global modeling capability. However, the continuous processing of full token sequences in deep Transformer layers introduces substantial inference overhead. Existing token pruning methods often struggle to balance efficiency and task adaptability: retraining-free methods rely on fixed importance heuristics, while retraining-based methods typically require additional optimization of the full model. Therefore, a key challenge is: how to learn token selection policies with both task awareness and long-term decision-making capability without updating large-scale pretrained ViTs. To address this challenge, we propose FASTER, a lightweight reinforcement learning(RL) framework for rapid task-adaptive ViT token pruning. FASTER freezes the pretrained ViT backbone and task head, and only optimizes a lightweight Actor–Critic policy network. It introduces Task-aware Sensitivity Encoding (TSE) to transform task-related knowledge embedded in the pretrained model into a reinforcement learning state representation, and learns instance-adaptive token retention policies by jointly considering token representations, spatial information, and budget constraints. Experimental results demonstrate that FASTER effectively reduces computational cost and inference latency on ViT-L, DeiT-B, DeiT-S, and DeiT-T, achieving a competitive trade-off between accuracy and

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.