QPrune: Quantum-Inspired Visual Token Pruning via Mixed-State Reduction on LVLMs
Abstract
Multimodal large language models (MLLMs) process a large number of visual tokens, resulting in substantial inference overhead, which has motivated extensive research on visual token pruning. Existing methods commonly identify visual tokens based on attention scores, visual saliency, feature similarity, or cross-modal relevance. However, identifying informative visual tokens requires accounting for the interactions between visual and textual tokens, which together form a complex information state. Quantum information theory offers an alternative perspective for analyzing such a complex information state. In this work, we propose QPrune, a quantum-inspired visual token pruning method centered on mixed-state representation. QPrune represents the visual token set as a mixed state, where each visual token is represented as a pure state with an associated probability under the current image-text pair. Based on the mixed state, state-aware token pruning considers both token relevance and state overlap with the retained tokens to reduce redundancy. State-preserving token refinement further supplements local visual information that is insufficiently represented after pruning. Extensive experiments demonstrate that QPrune effectively reduces the number of visual tokens while maintaining strong task performance across different pruning ratios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.