Language-Aware Efficient Adaptive Filtering for Visual Token Selection via Truncated Probing and Hierarchies
Abstract
Large vision-language models (VLMs) process hundreds or thousands of visual tokens, many of which are redundant for a given question. Existing training-free reduction methods often apply uniform transformations or rank patches independently using a single importance signal, making it difficult to preserve both spatial coverage and task-relevant detail under aggressive compression. We introduce LEAF, a training-free framework that jointly determines _how_ the image should be partitioned, _which_ patches are relevant, and _how much detail_ each region requires. LEAF combines a saliency-guided adaptive hierarchy with question-aware relevance extracted through truncated language-model probing, then allocates the token budget using a coverage-detail objective that balances regional coverage, patch relevance, and feature heterogeneity. Across four frozen VLMs and nine benchmarks, LEAF achieves the highest normalized average performance at all evaluated model-budget settings. On Qwen2.5-VL, it preserves 96.08% of full-token performance at 10% retention, compared with 90.76% for the strongest competing method, and achieves a BD-Token of -75.94% relative to random retention across the evaluated budget range, demonstrating effective visual-token reduction without training or backbone modification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.