acceptodds
Under review as a conference paper at ICLR 2027

Irreplaceability Matters: Counterfactual Intervention for Visual Token Pruning in MLLMs

Abstract

Multimodal Large Language Models (MLLMs) represent images as long visual-token sequences, making inference costly. Training-free pruning methods typically rank tokens by saliency or redundancy cues from a single forward pass, yet prominently scored tokens can remain compensable under tight budgets—importance alone does not ensure that visual evidence is hard to replace. We introduce CiVi, a plug-and-play inference-time selector that augments factual-view saliency and pairwise non-redundancy with a feature-space intervention-response proxy inspired by counterfactual reasoning. CiVi perturbs encoder features, compares clean and perturbed projected representations for ranking, and passes only retained clean tokens to the decoder. The intervention response provides a low-cost proxy for decoder-relevant sensitivity rather than an identified causal effect; decoder-level interventions and two-cue ablations support its decoder relevance and the complementarity of the three signals. Across LLaVA-, Qwen-, and video MLLM backbones, CiVi achieves a strong accuracy–efficiency trade-off, particularly under aggressive pruning. On LLaVA-1.5-7B, it preserves 95.7% of the unpruned nine-task average using only 64 of 576 visual tokens. Our code is available at https://anonymous.4open.science/r/CiVi-1082.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.