Output-Aligned Visual Token Routing for Efficient Vision-Language Models
Abstract
Visual token pruning accelerates Vision-Language Models (VLMs), yet existing methods predominantly rely on intermediate proxies such as attention weights or feature diversity to rank tokens. Visual elements typically exhibit strong spatial and semantic interdependencies. Consequently, ranking tokens in isolation fails to preserve the joint context required for generation, a challenge we term proxy metric misalignment. We propose Output-Aligned Routing (OAR), which supervises visual token selection directly via the inherent generation behavior of a frozen base VLM. In the offline stage, the base model uses unlabeled calibration data to evaluate retained visual subsets through the output KL divergence between unpruned and pruned generation distributions. We distill these divergence signals into soft inclusion targets to train a 1.91M-parameter Router. At inference time, the Router selects and gathers the top- tokens in a single, query-aware pass prior to language-model prefill, requiring zero online intervention from the base model. Across demanding benchmarks spanning charts, documents, infographics, and scene text, OAR preserves 92–95% of full-model accuracy on ChartQA and TextVQA while retaining only 80 visual tokens, an approximately 85% reduction in the visual sequence. Crucially, it achieves up to a 2.94 end-to-end speedup and a 4.8 reduction in KV cache memory, with an online selection overhead of under 1 ms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.