Beyond Semantic Matching: Let Modalities Talk through Their Position-Content Dual Manifolds for VLM Token Pruning
Abstract
Visual token pruning is an effective way to reduce the inference cost of VLMs, yet identifying a compact and sufficient subset of visual tokens remains challenging, as token importance depends not only on cross-modal relevance but also on complex, heterogeneous **Where–What** relations in each modality. Existing methods either prematurely mix position and content information into a shared representation or model them in separate stages, potentially obscuring their distinct structures and intrinsic dependency-relations. To address this, we introduce **Position-Content Dual Manifold (PCDM)**, a unified multimodal, multi-attribute relational system for visual token pruning. PCDM factorizes each token into coupled position and content nodes, thus forming two complementary manifolds per modality: a position manifold preserving native structural proximity and a content manifold capturing semantic correlations, with one-to-one node correspondence linking them together. The vision and language PCDMs are further connected through semantic alignment, enabling information to propagate globally across modalities and attributes while preserving their distinct structures. This enables inference of higher-order task relevance and structured pairwise affinities among visual tokens. Token selection is then formulated as a budget-constrained quadratic optimization that rewards task relevance while suppressing semantic/spatial redundancy. Experiments across *five VLMs and ten benchmarks* demonstrate consistently superior efficiency–accuracy trade-offs over state-of-the-art methods. On LLaVA-NeXT-7B, PCDM retains 94.5% of full performance with only 160 visual tokens (5.6%).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.