acceptodds
Under review as a conference paper at ICLR 2027

OTT-Vid: Optimal Transport Temporal Token Compression for Video-LLMs

Abstract

As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to visual tokens accumulated across frames. Training-free token compression reduces this overhead in part by exploiting repetition across frames, often using cross-frame similarity to identify redundant tokens or group similar frames for compression. However, cross-frame similarity and segmentation alone do not capture each token's importance within its frame or directly determine how much compression to apply to each frame pair. In this work, we propose OTT-Vid, a training-free outer-LLM compression framework that incorporates intra-frame token importance into temporal matching and derives frame-pair budgets from the resulting transport difficulty. After spatial pruning retains representative tokens, we formulate optimal transport (OT) between neighboring frames using importance-aware masses and a locality-aware cost combining feature dissimilarity with spatial distance. Smaller masses limit the coupling values of important tokens, reducing their priority in compression candidate ranking. The transport plan guides token-level compression, while its total cost serves as a proxy for compression difficulty, assigning larger compression budgets to lower-cost frame pairs. Experiments on six benchmarks spanning video question answering and temporal grounding across four backbones demonstrate strong performance at multiple token retention ratios, with gains in temporal grounding supporting the preservation of temporal evidence under compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.