Coverage-Preserving Recasting and Hierarchical Visual Token Pruning for Efficient Multimodel LLMs
Abstract
Multimodal Large Language Models increasingly rely on long visual token sequences, making token pruning essential for reducing inference overhead. However, existing methods often suffer from accuracy degradation under stringent token budgets since they insufficiently account for the spatial distribution of visual evidence and the layer-dependent characteristics of multimodal representations. Therefore, we investigate the mechanisms underlying effective token reduction and find that visual token importance in the encoder should be assessed through complementary visual cues together with crop-wise allocation to avoid overlooking informative regions, while instruction-conditioned relevance becomes reliable as visual semantics mature across decoder layers. Motivated by these findings, we propose Layered-Crop Holistic Token Pruning called LCHTP, a training-free two-stage framework that integrates spatially aware selection with layer-sensitive compression. We further develop a theoretical formulation that links pruning-induced representation distortion to prediction stability and inference complexity, which provides a principled account of the accuracy–efficiency trade-off. Extensive evaluations across diverse MLLMs and benchmarks demonstrate that LCHTP consistently preserves strong predictive performance while delivering inference acceleration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.