ConTra: Visual Token Pruning via Contrast Invariance in MLLMs
Abstract
Multimodal Large Language Models (MLLMs) incur substantial inference costs from the sheer volume of visual tokens, especially in text-rich scenarios where textual and visual semantics are deeply entangled within visual tokens. Existing pruning methods rank tokens by holistic visual salience and thus frequently discard text-critical tokens as redundant background. We address this limitation with the Contrast Invariance Principle—textual semantics remain stable under image-level contrast perturbations, whereas background semantics drift substantially. Because glyphs are only the extremal case of ordinal structure in images, the same principle transfers safely to natural, largely text-free scenes. Building on this principle, we propose ConTra, a training-free visual token pruning framework that generates a contrast-transformed view via the Chromatic Contrast Transformation (CCT) and computes cross-view token consistency to identify and preserve text-critical tokens. ConTra requires no attention matrices and only one batched vision-encoder pass over the two views, making it FlashAttention-compatible and readily deployable on off-the-shelf MLLMs without retraining. Across 22 benchmarks spanning multiple model scales, ConTra retains 90.1% performance at a 70% pruning rate on text-rich tasks—surpassing the strongest prior method by 8.1 points—while remaining fully competitive with query-guided methods on general tasks (98.1% retention) despite using no query information, and reduces FLOPs by up to 63%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.