AVTC: Aggressive Visual Token Compression via Set-Aware Selection for Video Large Language Models
Abstract
Visual token compression offers a practical way to reduce the inference cost of video large language models (Video-LLMs), but preserving model performance under aggressive compression remains challenging. Existing methods often select tokens by individual importance or prune them based on pairwise similarity, which can lead to redundant selections and incomplete coverage of the original visual content under tight budgets. To address this limitation, we introduce AVTC, a training-free framework that performs set-aware selection by evaluating how well the selected tokens jointly reconstruct the original visual features. Specifically, AVTC scores each candidate by how much its addition reduces the total importance-weighted projection error of all original tokens onto the subspace spanned by the selected anchors. This criterion favors tokens that contribute important information not yet captured by the current selection. Experiments on three Video-LLMs and four benchmarks show that AVTC achieves the highest average accuracy among the compared methods at layer-average token retention ratios of 2%, 5%, and 10%. On LLaVA-OneVision-7B, AVTC retains 94.6% of the uncompressed model’s average accuracy at just 2% token retention. At 10% retention, it achieves 100.2% of the uncompressed model’s average accuracy while delivering a 6.6 prefill speedup and a 2.0 speedup in time to first token (TTFT). Code is available at https://anonymous.4open.science/r/AVTC-8F87.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.