Complete the Scene, Cover the Video: Recurrent Token Compression for Video LLMs
Abstract
Visual token compression aims to reduce inference costs in Video Large Language Models while preserving information for video understanding. Existing compressors often build on independent frame-wise selection of important tokens. However, such selection may repeatedly retain content already well represented in earlier outputs while overlooking information not yet preserved, limiting coverage of content unfolding throughout the video. Motivated by representation gaps revealed in our video-level feature coverage analysis, we propose C²V, a training-free, before-LLM method using the preceding frame’s compressed representation as a visual state. It greedily selects anchors that complement the state in representing the current scene. Candidate sets are evaluated by Maximum Mean Discrepancy (MMD) between the attention-weighted scene distribution and an optimal convex mixture of the state and candidate anchor distributions. Mass-preserving aggregation incorporates unselected tokens’ features and probability mass into the anchors, producing LLM input features and the anchor distribution used as the next visual state. We evaluate C²V on five benchmarks spanning short clips to long videos, using Qwen3-VL-8B, InternVideo3-8B, and Qwen2.5-VL-7B at four retention ratios (10–25%). C²V achieves the highest five-benchmark average accuracy among evaluated compressors in all 12 model–retention settings. It also outperforms competing methods in video-level feature coverage at matched token budgets, indicating better collective representation of source features.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.