Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Visual Token Pruning in MLLMs
Abstract
Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring substantial inference costs. While existing visual token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analysis reveals that ranking pruning methods by average benchmark accuracy conceals substantial sample-wise complementarity: although the average-best strategy excels overall, alternative strategies prove superior on a considerable fraction of individual samples. To harness this diversity, we propose **VIP-Router**, a lightweight **VI**sion **P**runing **Router** that, conditioned on low-cost visual and textual features, selects for each input the pruning strategy predicted to be most suitable at a specified pruning level, while retaining full-token inference as a fallback when pruning is predicted to be unfavorable. On VTC-Bench Group A, a pruning-sensitive suite spanning eight vision-language benchmarks, VIP-Router consistently outperforms the best fixed pruning strategy at every reduction ratio, with relative gains of 26.9% in average accuracy and 22.0% in average utility, a cost-aware metric that penalizes each method for the visual tokens it actually retains. Under stricter image-disjoint splits, VIP-Router still improves average utility over the best fixed strategy by 4.7% in relative terms, with gains in all 15 runs (three splits × five seeds). Moreover, VIP-Router is plug-and-play: it modifies neither the pruning algorithms nor the model weights and adds only 1.38M trainable parameters (0.017% of the backbone). It also remains effective on four additional MLLMs and yields gains on four unseen benchmarks, highlighting the potential of sample-adaptive routing for visual token pruning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.