DAPrune: Dual-Adaptive Pruning via Adaptive Ratio and Layer Selection for Large Vision Language Models
Abstract
Visual tokens in large vision language models (LVLMs) increase rapidly with image resolution and video inputs, resulting in substantial computational overhead. To reduce this cost, a pruning policy is needed to reduce the number of tokens involved in the attention computation. However, existing token pruning methods mostly adopt fixed pruning policies for all inputs, which are not sufficiently efficient and cannot adapt to input difficulty. To address this limitation, we propose DAPrune, a training-based dual-adaptive pruning framework that jointly learns where to prune and how many visual tokens to retain for each input. Specifically, DAPrune introduces two lightweight modules: a routing module for pruning layer selection and a ratio module for adaptive token retention. With the base LVLM frozen, these modules can adapt efficiently within 3 hours on a single H100 GPU for LLaVA-1.5. Empirically, DAPrune preserves over 98.5% of the baseline performance under high FLOPs reduction on LLaVA-1.5, LLaVA-NeXT, and Qwen-2.5-VL, consistently outperforming prior methods under comparable settings. Furthermore, theoretical analysis suggests that, under the same average budget, our adaptive policy achieves lower representation error than fixed pruning strategies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.