acceptodds
Under review as a conference paper at ICLR 2027

Language-Aware Efficient Adaptive Filtering for Visual Token Selection via Truncated Probing and Hierarchies

Abstract

Large vision-language models (VLMs) process hundreds or thousands of visual tokens, many of which are redundant for a given question. Existing training-free reduction methods often apply uniform transformations or rank patches independently using a single importance signal, making it difficult to preserve both spatial coverage and task-relevant detail under aggressive compression. We introduce LEAF, a training-free framework that jointly determines _how_ the image should be partitioned, _which_ patches are relevant, and _how much detail_ each region requires. LEAF combines a saliency-guided adaptive hierarchy with question-aware relevance extracted through truncated language-model probing, then allocates the token budget using a coverage-detail objective that balances regional coverage, patch relevance, and feature heterogeneity. Across four frozen VLMs and nine benchmarks, LEAF achieves the highest normalized average performance at all evaluated model-budget settings. On Qwen2.5-VL, it preserves 96.08% of full-token performance at 10% retention, compared with 90.76% for the strongest competing method, and achieves a BD-Token of -75.94% relative to random retention across the evaluated budget range, demonstrating effective visual-token reduction without training or backbone modification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.