ALIGN AND BRIDGE: VISUAL TOKEN PRUNING IN THE PRESENCE OF A MODALITY GAP
Abstract
Vision-language models (VLMs) achieve strong performance across multimodal tasks, but their numerous visual tokens incur substantial computational and memory costs. Pruning visual tokens before the LLM decoder reduces the input sequence length, yet identifying query-relevant tokens remains challenging in the presence of a modality gap and limited cross-modal interaction. We propose Align and Bridge, a framework that learns query-conditioned relevance before the decoder and propagates it throughout progressive token pruning. Its nonlinear Vision-Text Aligner (VTA) is trained solely on pre-decoder visual and textual embeddings without executing the LLM decoder. Score-dispersion regularization encourages clearer separation between query-relevant regions and background. Within the decoder, Prior-aware Diverse Relevance (PDR) progressively prunes visual tokens by combining text-to-vision attention with pre-decoder importance priors and token diversity. This design connects early token selection with subsequent pruning to preserve relevant visual information while suppressing redundancy. Experiments across multiple VLMs and multimodal benchmarks show improved performance retention on most evaluated backbones, with substantial reductions in visual tokens and FLOPs; on Qwen2.5-VL, a learned baseline remains stronger. On Qwen3-VL-8B, at a nominal budget of 22.2% of visual tokens averaged across decoder layers, our method preserves 93.4% of the original QA performance and improves relative visual grounding performance by about 12.0 percentage points over the strongest grounding baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.