RDPrune: Learning Visual Token Pruning via Rate-Distortion Optimization for Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) incur substantial inference overhead due to long visual token sequences. Existing visual token pruning methods typically operate under predefined budgets, overlooking that inputs of different difficulty may require different amounts of visual information. We formulate visual token pruning as lossy compression, where the retained token ratio defines the rate and the behavioral discrepancy induced by pruning defines the distortion. Joint rate-distortion optimization over the data distribution enables adaptive allocation of a shared token budget across inputs according to their difficulty. We instantiate this formulation with RDPrune, a lightweight coarse-to-fine framework consisting of an early selector after vision encoder and a mid-layer selector inside the LLM. Both selectors learn binary keep-or-drop decisions to progressively compress the visual tokens. Crucially, we explicitly model pruning-induced distortion at three complementary levels: task distortion via cross-entropy, prediction distortion via output distribution KL divergence, and representation distortion via hidden states discrepancy. This directly quantifies how token removal affects downstream model behavior, rather than relying solely on pre-pruning token importance proxies. Across seven benchmarks and three LVLM backbones, RDPrune retains 97.7% of the original performance with only 11.1% of the visual tokens, achieving 2.8 inference speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.