acceptodds
Under review as a conference paper at ICLR 2027

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

Abstract

Structured pruning reduces the parameter footprint of vision-language models (VLMs), but its effectiveness depends on how multimodal inputs are used to estimate component importance. Visual tokens often outnumber text tokens, and including redundant visual content in calibration can impair pruning decisions. We propose SlimVLM, a structured pruning framework that combines adaptive visual token selection with sensitivity-aware allocation of pruning ratios. At each decoder layer selected for pruning, SlimVLM uses average text-to-visual attention to select visual tokens for importance estimation while retaining all text tokens. It then adjusts the pruning ratio of each attention or MLP module according to the correlation between its pruned and original outputs, removes attention heads and MLP channels, and compensates for output deviations through least-squares fitting. Token selection operates during calibration; the deployed model uses the reduced parameter structure. Experiments on LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL demonstrate the effectiveness of this approach across eight multimodal benchmarks. At a 20% pruning ratio on LLaVA-1.5-7B, SlimVLM retains 93.80% of the unpruned model's average relative performance and leads the evaluated structured-pruning baselines on eight of nine reported task settings. Ablations further show that selecting visual tokens for calibration improves post-pruning performance in both SlimVLM and FLAP.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.