acceptodds
Under review as a conference paper at ICLR 2027

Do We Still Need Token Compression in Video Large Language Models?

Abstract

Visual-token pruning is widely used to improve the efficiency of video large language models (VideoLLMs). However, backbones with native dynamic-resolution support introduce an alternative: controlling the visual-token budget through input resolution before vision encoding. We revisit this design choice through a controlled study of Qwen3-VL-4B, GLM-4.6V-Flash, VideoLLaMA3, and Oryx, covering 64- and 128-frame inputs and four video question-answering benchmarks. Under matched actual token budgets and controlled video interfaces, established pruning methods do not consistently outperform random dropping, and method rankings vary substantially across backbones. Simple resolution reduction can also match or exceed the accuracy of post-encoding pruning. To understand these results, we compare random dropping with VCAST and examine the roles of spatiotemporal coverage, evidence aggregation, and backbone characteristics. Our analysis suggests that selecting informative tokens alone is insufficient: compression must also preserve the distribution and structure of visual evidence. Random dropping can retain broad spatiotemporal coverage, while downsampling aggregates local evidence before the vision encoder constructs a dense representation. End-to-end profiling further shows that resolution reduction saves vision-encoding computation that post-encoding pruning has already incurred. Building on these findings, we investigate dynamic resolution beyond uniform downsampling, separating budget fidelity from content-aware allocation through a budget-controlled oracle-to-heuristic study. Our results motivate dynamic resolution as a promising paradigm for efficient VideoLLMs: it directly reduces vision-encoding cost and offers a route to allocating visual computation adaptively across frames and task demands. More broadly, effective visual compression requires joint consideration of the backbone, evidence preservation, and where computation is reduced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.