DOVE: D-Optimal Visual Evidence for Extreme Video Token Compression
Abstract
Video large language models process long videos as dense sequences of visual tokens, making multimodal prefilling increasingly expensive. At extreme compression ratios, preserving task-relevant information becomes particularly challenging because decisive visual evidence can be sparse, transient, and distributed across different spatial and temporal contexts. We introduce DOVE (D-Optimal Visual Evidence), a training-free framework for extreme video-token compression. DOVE constructs a quality-weighted evidence geometry that combines visual semantics with spatial, temporal, dynamic, and optional query-dependent cues, and retains a fixed-size subset by maximizing the log-determinant of a ridge-regularized information matrix. Under this criterion, candidates aligned with already well-covered evidence directions provide limited additional volume, whereas reliable candidates spanning weakly covered directions yield larger marginal gains. We optimize the objective efficiently over a budget-proportional candidate pool using exact leverage-based greedy updates and positive-gain local refinement. A fixed-length residual readout recovers compatible discarded information without changing the selected set or output length. Across VideoMME, EgoSchema, LongVideoBench, and MVBench with three Video LLM backbones, DOVE outperforms the strongest baseline's average score by 3.5 points at 1% layer-averaged retention and by 2.9 points at 0.5% percentage. These results demonstrate the effectiveness of quality-weighted D-optimal evidence compression under extremely constrained visual-token budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.