Do We Still Need Token Compression in Video Large Language Models?
Abstract
Visual-token pruning is widely used to improve the efficiency of video large language models (VideoLLMs). However, backbones with native dynamic-resolution support introduce an alternative: controlling the visual-token budget through input resolution before vision encoding. We revisit this design choice through a controlled study of Qwen3-VL-4B, GLM-4.6V-Flash, VideoLLaMA3, and Oryx, covering 64- and 128-frame inputs and four video question-answering benchmarks. Under matched actual token budgets and controlled video interfaces, established pruning methods do not consistently outperform random dropping, and method rankings vary substantially across backbones. Simple resolution reduction can also match or exceed the accuracy of post-encoding pruning. To understand these results, we compare random dropping with VCAST and examine the roles of spatiotemporal coverage, evidence aggregation, and backbone characteristics. Our analysis suggests that selecting informative tokens alone is insufficient: compression must also preserve the distribution and structure of visual evidence. Random dropping can retain broad spatiotemporal coverage, while downsampling aggregates local evidence before the vision encoder constructs a dense representation. End-to-end profiling further shows that resolution reduction saves vision-encoding computation that post-encoding pruning has already incurred. Building on these findings, we investigate dynamic resolution beyond uniform downsampling, separating budget fidelity from content-aware allocation through a budget-controlled oracle-to-heuristic study. Our results motivate dynamic resolution as a promising paradigm for efficient VideoLLMs: it directly reduces vision-encoding cost and offers a route to allocating visual computation adaptively across frames and task demands. More broadly, effective visual compression requires joint consideration of the backbone, evidence preservation, and where computation is reduced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.