VOLT: Joint Token Selection and Frame Budget Allocation via Subspace Volume Maximization
Abstract
Efficient video understanding with multimodal large language models requires compressing visual tokens without losing sparse events or local details. Existing training-free approaches commonly allocate equal token budgets to frames or separate frame-level budget prediction from token selection. These designs can retain redundant content and underrepresent frames containing informative events. We present VOLT, which unifies token selection and cross-frame budget allocation as constrained subspace volume maximization. A shared log-determinant objective measures the geometric diversity of retained token representations and jointly determines their identities and distribution across frames. A minimum retention constraint preserves coverage of each frame, while within-frame nearest-neighbor merging transfers information from discarded tokens to retained tokens without increasing sequence length. Experiments on five video understanding benchmarks with three video MLLM backbones show an improved accuracy-efficiency trade-off under aggressive compression. VOLT requires neither retraining nor modification of the backbone, providing a practical geometric approach to video token reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.