CoGCo: Codec Guided Training-free Video-LLM Token Compressor
Abstract
Video Large Language Models (Video LLMs) process thousands of visual tokens from sampled frames, making redundant video representations a major source of inference cost. Training-free video token compression must preserve persistent scene context and subtle temporal changes, yet token saliency alone does not reveal which information requires an additional representative. We present CoGCo, which exploits the structure of standard video codecs to separate these complementary responsibilities. Reference-frame tokens establish semantic anchors through attention- and diversity-aware selection, supplemented by reference locations supported by subsequent updates. Prediction-frame selection combines camera-compensated motion vectors, residuals, and within-frame diversity to preserve complementary temporal evidence. It excludes a small fraction of the highest-attention tokens from independent selection, as our analysis suggests that such features primarily capture overall semantic context rather than fine-grained updates. Moreover, we leverage codec frame packets to efficiently allocate the token budget. Across four benchmarks and four backbones, CoGCo achieves higher aggregate accuracy than the SOTA training-free compression baselines at evaluated retention settings of 15%, 20%, and 25%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.