EviQ: Hierarchical Query-Conditioned Evidence Selection for Efficient VideoLLMs
Abstract
Video large language models (VideoLLMs) have achieved strong performance in video understanding, yet their large visual token budgets introduce substantial inference latency and memory consumption. Most existing training-free token compression methods rely primarily on visual cues, potentially overlooking query-critical evidence scattered across sparsely sampled frames, while some query-aware approaches incur additional computation through within-LLM attention or auxiliary models. Instead, we reveal that cosine similarity between projected visual tokens and embedded query tokens provides a lightweight relevance signal both across and within sparsely sampled frames. Building on this observation, we propose **EviQ**, a training-free pre-LLM token pruning framework that constructs a compact query-conditioned evidence set in two stages. **EviQ** first routes the global token budget across frames according to frame-level query relevance, and then selects tokens within each frame by jointly preserving informative visual context and promoting complementary coverage of query semantics. Extensive experiments on diverse VideoLLMs and multiple benchmarks demonstrate that **EviQ** achieves state-of-the-art accuracy-efficiency trade-offs among training-free token compression methods. Notably, **EviQ** preserves **99.9%** of vanilla LLaVA-OV-7B accuracy using only 25% of visual tokens, while achieving a 1.59 speedup and reducing GPU memory consumption by 8.7%. Code is released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.