acceptodds
Under review as a conference paper at ICLR 2027

CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models

Abstract

Multimodal Large Language Models (MLLMs) suffer from substantial computational overheads, driven by the massive redundancy of visual token sequences. Existing methods reduce this cost by pruning redundant visual tokens, typically using features from a single or fixed set of ViT layers together with a predefined token selection policy. However, adapting visual feature extraction and token selection to the varying information needs of different instructions remains challenging under these static designs. To address these limitations, we present CLass-Adaptive Layer fusion and dual-Stage Pruning (CLASP), a plug-and-play token reduction framework that jointly adapts visual feature fusion and token pruning to each instruction. Specifically, we construct a category-specific visual representation to balance fine-grained visual detail and high-level semantics. Then we perform dual-stage pruning that allocates the token budget between attention-salient pivots (relevance) and redundancy-aware completion tokens (coverage). Mechanism analyses further clarify how category-conditioned fusion and relevance–coverage token selection support instruction-adaptive pruning. Extensive evaluations across diverse benchmarks, pruning ratios, and MLLM architectures demonstrate a favorable balance between accuracy and efficiency. Taken together, these results show that instruction-adaptive visual processing can reconcile aggressive token reduction with robust multimodal performance, offering a practical path toward MLLM deployment under constrained computational budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.