QUEST: Query-guided Evidence Selection for Visual Token reduction
Abstract
Long visual token sequences increase the inference cost of the large language model (LLM) in vision-language models (VLMs). Effective visual token reduction requires preserving the visual information needed to answer the textual query. An external vision-language alignment model (VLAM) can provide relevance between image regions and the query. However, directly mapping this relevance to the visual tokens of the target VLM is difficult because the VLAM and the target VLM may use different vision encoders and visual representations. To address this issue, we propose QUEST, a training-free Pre-LLM visual token selection method. QUEST uses the relevance from an external VLAM to identify query-relevant Evidence Seed units and uses the corresponding vision encoder tokens as Evidence Seeds. It constructs Evidence Attention from the attention between the Evidence Seeds and other visual tokens at an intermediate layer of the target VLM's vision encoder. QUEST also selects an Evidence Source Layer based on layer-wise attention distance. Evidence Attention is combined with Visual Guidance, followed by diversity-aware selection to choose the final visual tokens under a given token budget. We evaluate QUEST on LLaVA-NeXT-7B and Qwen2.5-VL-7B across multiple VLAMs, token budgets, and nine benchmarks. The results show high average performance retention relative to the full-token VLMs across the evaluated token budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.