acceptodds
Under review as a conference paper at ICLR 2027

CompVID: Which Video Regions Deserve More Tokens? Complexity-Aware Token Compression for Video LLMs

Abstract

Existing Video Large Language Model (VLLM) token pruning methods typically perform importance- or diversity-based visual token compression at the video, segment, or frame level. However, they overlook an inherent property of video data: video complexity is unevenly distributed across temporal and spatial dimensions, leading to suboptimal token allocation across spatiotemporal regions. Through empirical analysis, we find that allocating tokens to higher-complexity regions yields larger logit margin gains and contributes to improving video understanding capabilities. Inspired by this, we introduce a new perspective: regional complexity determines how much representation budget a region requires, whereas dimension-wise complexity determines how the region should be partitioned. Building on this insight, we propose CompVID, a training-free framework for complexity-aware video token allocation. Specifically, CompVID employs dimension-wise singular value decomposition (SVD) to reveal the rank structure along each dimension and quantify its intrinsic complexity through singular-value decay. Based on these complexity estimates, we introduce SVDtree, an SVD-guided adaptive partition tree that recursively refines the video feature volume along the complexity-dominant dimension and allocates token budgets to the resulting regions according to their regional complexity. Evaluations on VLLMs across video understanding benchmarks demonstrate that CompVID achieves a state-of-the-art efficiency–accuracy trade-off. By pruning of visual tokens on LLaVA-OneVision-7B, CompVID reduces TFLOPs by and accelerates the LLM prefill stage by , while preserving of the full-token performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.