acceptodds
Under review as a conference paper at ICLR 2027

Where to Look, What to Keep: Two-Phase Learning for Grounding-Aware Visual Token Selection

Abstract

Multimodal Large Language Models (MLLMs) encode images into hundreds or thousands of visual tokens, incurring substantial inference overhead. Although visual token pruning mitigates this cost, aggressive pruning can impair visual grounding, which requires both recognizing the referred target and predicting its spatial coordinates. We diagnose this degradation with budget-matched token-addition experiments and find that existing pruners either lose focus on the target or discard the spatial context needed for grounding. These findings motivate two complementary learning objectives: learning query-conditioned target relevance (where to look) and optimizing the retained subset for grounding accuracy (what to keep). We implement these objectives through a two-phase framework that trains a lightweight, query-conditioned token scorer on frozen MLLM backbones. In Phase One, Localization Prior Learning (LPL) uses bounding-box supervision to teach the scorer where to look. In Phase Two, Grounding-Guided Selection Refinement (GSR) optimizes token selections through direct grounding feedback, teaching the scorer what to keep by rewarding subsets that yield accurate bounding-box predictions. At an visual-token retention ratio, our method retains of unpruned grounding accuracy on Qwen2.5-VL-7B and exceeds the unpruned LLaVA-1.5-7B (), while retaining and of general VQA performance across ten benchmarks, respectively. Furthermore, our approach achieves up to end-to-end speedup and prefill speedup, offering a practical solution for efficient MLLM inference. Code and model checkpoints will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.