Rethinking Post-Pruned Evidence: Counterfactual Decision Reasoning for Fixed-Budget Visual Token Refinement
Abstract
Visual token pruning improves the efficiency of vision-language large models (VLLMs) but often discards query-relevant evidence under aggressive compression, leading to degraded model decisions. Existing post-pruning recovery methods restore discarded information according to predefined relevance or representation criteria, without explicitly evaluating how the resulting token composition affects the current decision. Consequently, restoration introduces a recovery–risk trade-off: useful evidence is recovered, while unfavorable interventions disrupt otherwise reliable predictions. To address this issue, we propose Counterfactual-Aware Restoration through Token Exchange (CARE), a training-free and fixed-budget post-pruning recovery framework based on query-adaptive token exchange and counterfactual decision-guided selection. Ours first identifies replaceable retained tokens and retrieves complementary discarded evidence to construct diverse one-to-one exchange proposals without increasing the visual-token budget. Ours then treats alternative token compositions as counterfactual interventions and evaluates their decision-level effects against the original pruned representation using the frozen VLLM. An alternative composition is adopted only when it provides stronger decision support than the no-change reference. Experiments across multiple benchmarks, pruning methods, and token budgets demonstrate consistent performance improvements, substantially fewer recovery-induced errors, and preserved pruning efficiency without additional training or token-count increase.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.