acceptodds
Under review as a conference paper at ICLR 2027

VOLT: Joint Token Selection and Frame Budget Allocation via Subspace Volume Maximization

Abstract

Efficient video understanding with multimodal large language models requires compressing visual tokens without losing sparse events or local details. Existing training-free approaches commonly allocate equal token budgets to frames or separate frame-level budget prediction from token selection. These designs can retain redundant content and underrepresent frames containing informative events. We present VOLT, which unifies token selection and cross-frame budget allocation as constrained subspace volume maximization. A shared log-determinant objective measures the geometric diversity of retained token representations and jointly determines their identities and distribution across frames. A minimum retention constraint preserves coverage of each frame, while within-frame nearest-neighbor merging transfers information from discarded tokens to retained tokens without increasing sequence length. Experiments on five video understanding benchmarks with three video MLLM backbones show an improved accuracy-efficiency trade-off under aggressive compression. VOLT requires neither retraining nor modification of the backbone, providing a practical geometric approach to video token reduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.