Retained but Underused: Modality Attention Imbalance after Visual Token Pruning
Abstract
Visual token pruning accelerates vision-language models (VLMs) by retaining only a small yet informative subset of image tokens. Most existing methods focus on token selection and assume that the decoder can effectively use the retained visual evidence. However, we find that pruning also changes the attention allocation between visual and textual modalities. As the token budget decreases, the average attention received by each retained visual token increases, while aggregate visual attention relative to textual attention decreases. We refer to this phenomenon as modality attention imbalance. Meanwhile, task accuracy and teacher-forced next-token prediction loss may even vary in the same direction. The output distribution also shifts toward the distribution produced without an image, indicating greater reliance on textual context. Simply rebalancing attention across all decoder layers fails to correct this imbalance. Through controlled causal interventions, we identify task-specific decoder regions where visual evidence is critical. These regions vary across models and tasks but transfer across pruning methods. Based on these findings, we propose Causal Attention Recalibration (CAR), a training-free post-pruning framework consisting of Causal Attention Profiling and Plateau Attention Recalibration. Causal Attention Profiling identifies decoder regions where visual evidence is causally important. Plateau Attention Recalibration applies a flat-top profile within these regions to rebalance visual and textual attention. Experiments across multiple VLMs, benchmarks, and representative pruning methods demonstrate performance gains of up to 2.38% with negligible computational overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.