Bifocal: Training-Free Mixed Granularity Visual Token Pruning for Multimodal LLMs
Abstract
Training-free visual token pruning allocates a fixed budget among candidate tokens, and existing methods keep refining the scoring function, hovering near the ceiling of the general benchmarks, yet most collapse on fine-grained tasks such as grounding. We instead ask what visual evidence the LLM needs at inference, and find two answers. First, the evidence can come from only two sources, the image and the question, and no single score serves every task. Second, keeping only a subset of native patches breaks the integrity of the visual sequence in both content and position. Building on both findings, we present **Bifocal**, a training-free, two-stage visual token pruning method built on mixed-granularity evidence. Before the LLM, it keeps two kinds of fine evidence and a coarse context. Query-focused evidence keeps the tokens the question refers to, image-intrinsic evidence maximises how well the token set represents and covers the image, and a pooled coarse context preserves the dual integrity of the retained tokens. Inside the LLM, it further keeps the top-K tokens by text-to-image attention. Evaluated across three models and several budgets, spanning general VQA, reasoning, OCR, hallucination and grounding, Bifocal achieves state-of-the-art results: it attains the best general average in every setting while, at the tightest budget, leading the strongest baseline on grounding by to percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.