CausalPrune: Counterfactual Causal Token Pruning for Efficient Large Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, but their practical deployment is still limited by the substantial computation introduced by excessive visual tokens. Visual token pruning offers an effective solution by removing redundant visual tokens during inference. However, existing methods typically estimate token importance from correlational statistics, which cannot distinguish tokens that causally determine the model output from those spuriously associated with it. To address this issue, we propose CausalPrune, a counterfactual causal token pruning framework for efficient LVLMs. We formulate LVLM inference as a Structural Causal Model and introduce the Causal Token Importance Score (CTIS), which estimates the Average Treatment Effect of intervening on each visual token. To make causal estimation practical, we derive an efficient first-order approximation of CTIS and incorporate backdoor adjustment to mitigate latent cross-modal confounding. Within this causal formulation, layer-wise causal mediation decomposes token importance into direct and indirect effects through intermediate hidden states, enabling progressive pruning across transformer layers in a principled manner. Extensive experiments on various LVLM architectures, such as LLaVA-1.5 and Qwen2.5-VL, demonstrate that CausalPrune maintains comparable or even superior multimodal performance while achieving substantial reductions in FLOPs and prefill latency, consistently outperforming existing pruning methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.