acceptodds
Under review as a conference paper at ICLR 2027

DynVid: Adaptive Dynamism-Aware Token Compression for Efficient Video LLMs

Abstract

Video Large Language Models (VLLMs) have achieved remarkable success in video understanding, yet suffer from prohibitive computational cost due to the massive number of visual tokens from densely sampled frames. While visual token compression offers a promising solution, existing methods are limited in two respects: they either allocate a uniform token budget across all frames, neglecting the inherent variation in scene complexity, or primarily rely on attention scores for token selection, overlooking temporal dynamics unique to video. To address these limitations, we propose DynVid, a training-free, dynamism-aware adaptive token compression framework for VLLMs that follows an allocate-then-select paradigm. In allocation stage, video frames are partitioned into semantically coherent segments via dynamic segmentation, and each segment receives an adaptive token budget proportional to its visual complexity. In selection stage, we introduce a token-level dynamism score as a temporally-aware importance signal that complements attention scores, and employ a Facility Location objective to select anchor tokens that are both individually important and collectively representative. Extensive experiments on eight video understanding benchmarks demonstrate that DynVid outperforms training-free baselines in both accuracy and inference efficiency. Notably, on LLaVA-OneVision, DynVid retains 99.3% relative accuracy with 10% average retention ratio of visual tokens preserved, and still maintains 96.4% relative accuracy at an aggressive 5% retention ratio. On Qwen2.5-VL, DynVid achieves 97.4% and 96.0% relative accuracy at 10% and 7.5% retention ratios, respectively, while delivering up to 7.53 and 9.49 prefilling speedups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.