acceptodds
Under review as a conference paper at ICLR 2027

Output-Aligned Visual Token Routing for Efficient Vision-Language Models

Abstract

Visual token pruning accelerates Vision-Language Models (VLMs), yet existing methods predominantly rely on intermediate proxies such as attention weights or feature diversity to rank tokens. Visual elements typically exhibit strong spatial and semantic interdependencies. Consequently, ranking tokens in isolation fails to preserve the joint context required for generation, a challenge we term proxy metric misalignment. We propose Output-Aligned Routing (OAR), which supervises visual token selection directly via the inherent generation behavior of a frozen base VLM. In the offline stage, the base model uses unlabeled calibration data to evaluate retained visual subsets through the output KL divergence between unpruned and pruned generation distributions. We distill these divergence signals into soft inclusion targets to train a 1.91M-parameter Router. At inference time, the Router selects and gathers the top- tokens in a single, query-aware pass prior to language-model prefill, requiring zero online intervention from the base model. Across demanding benchmarks spanning charts, documents, infographics, and scene text, OAR preserves 92–95% of full-model accuracy on ChartQA and TextVQA while retaining only 80 visual tokens, an approximately 85% reduction in the visual sequence. Crucially, it achieves up to a 2.94 end-to-end speedup and a 4.8 reduction in KV cache memory, with an online selection overhead of under 1 ms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.