GLANCE: Efficient Video Large Language Models via Screening and Parallel Redundancy Suppression
Abstract
Denser frame sampling improves temporal coverage in video large language models (Video-LLMs) but makes decoder prefill expensive. Visual-token compression must balance subset quality with selection cost: token-wise ranking can retain redundant evidence, whereas subset-aware selection incurs pairwise comparisons and sequential updates. We introduce , a training-free, query-agnostic video token compressor that compares redundancy within compact candidate pools. Informativeness and Uniqueness-guided Screening (IUS) combines within-group feature-response ranks with video-wide directional uniqueness to shortlist candidates. Parallel Redundancy Suppression (PRS) discounts similarity to earlier candidates in a fixed utility order, computing penalties in parallel without iterative selected-set updates. Across three Video-LLM backbones and four benchmarks, the complete configuration, with approximately 12.5% decoder-input retention and late visual-token withdrawal at nominal 10% retention, preserves 94.2–97.6% of full-token mean accuracy with – speedups in model-side time to first token (TTFT). In a matched Qwen3-VL comparison, PRS reduces candidate-selection time from 8.37 to 0.66 ms at 32 frames with close observed accuracy. Compressed 128-frame input also exceeds the 64-frame full-token baseline in three-task mean accuracy at lower TTFT on Qwen3-VL.Code is available at https://anonymous.4open.science/r/glance-CEE3
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.