ReCo:Training-Free Visual Token Pruning via Relational Transition Coreset in Large Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) incur substantial inference overhead and KV cache costs because visual inputs are represented as long, dense token sequences. Among existing pruning methods, attention-based approaches compress high-dimensional interactions into scalar saliency scores that tend to exhibit instability across layers and attention heads, whereas redundancy elimination methods often equate unimodal feature similarity with cross-modal semantic substitutability, risking the loss of fine-grained visual evidence. Both paradigms typically evaluate tokens within an isolated representation space either prior to or following cross-modal projection, failing to utilize the information conveyed by pairwise relational transitions across the projection. We propose ReCo, a training-free visual token pruning framework based on relational transition coresets. ReCo jointly characterizes pre-projection visual relations, post-projection language-aligned relations, and projection-induced relational shifts, integrating them into a unified relational transition kernel. Visual token pruning is subsequently formulated as a submodular facility-location problem, where a greedy selector constructs a compact coreset by maximizing marginal relational coverage. Experiments on LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL demonstrate that ReCo achieves competitive performance across both fixed-length and dynamic-resolution inputs. It maintains 91.02% of the performance on Qwen2.5-VL even under an aggressive 90% visual token reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.