Reroute, Don’t Remove: Recoverable Visual Token Routing for Vision-Language Models
Abstract
Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory. Many decoder-side visual-token pruning methods follow a rank-and-remove paradigm: they score visual tokens, keep a compact subset, and permanently discard the rest. We show that this irreversible action can be fragile because visual-token importance changes across decoder depth; tokens ranked low at one stage may become relevant in later layers, especially for grounding- sensitive queries. We propose Reroute, a training-free, scorer-preserving plug-in that replaces removal with recoverable routing. At each routing stage, selected visual tokens pass through decoder blocks, while deferred tokens bypass the stage and re-enter the candidate pool at the next routing decision. By preserving the underlying ranking signal and modifying only the post-selection action, Reroute enables a controlled study of irreversible visual-token pruning. Reroute inherits existing attention-score ranking rules within a stage-wise routing framework, largely preserving the TFLOPs and KV-cache efficiency of the pruning method it augments while retaining most of the corresponding speedup. Across LLaVA-1.5 and Qwen backbones, Reroute improves matched-budget grounding performance across decoder-side visual-token pruning settings while maintaining general VQA performance under aggressive token reduction. These results suggest that VLM token reduction should not be viewed only as irreversible pruning, but also as recoverable routing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.