acceptodds
Under review as a conference paper at ICLR 2027

SCOUT: Visual Token Pruning for Multimodal Large Language Models via Semantic and Cross-modal Guidance

Abstract

Multimodal Large Language Models (MLLMs) have achieved strong performance in vision–language understanding by integrating vision encoders with large language models. However, the quadratic complexity of attention with respect to token length makes inference expensive, with visual tokens dominating the input sequence. Recent studies have explored visual token pruning to address this issue, but most methods estimate token importance solely based on attention scores. In certain cases, such scores may not reliably reflect token importance for two reasons: attention in the vision encoder arises from unimodal self-attention, while attention concentration can assign high weights to regions with limited semantic relevance. To address these limitations, we propose SCOUT (Semantic and Cross-modal guided visual Token pruning), a novel two-stage visual token pruning framework. In Stage 1, SCOUT combines encoder attention with value norms to mitigate the limitations of relying solely on attention scores and performs semantic region-aware coarse filtering. In Stage 2, it uses text-to-visual attention for fine-grained pruning within the language model, with entropy-guided gating to reduce the influence of attention sink-like behavior. Experiments on multiple MLLM architectures across image and video benchmarks show that SCOUT consistently improves the trade-off between efficiency and performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.