acceptodds
Under review as a conference paper at ICLR 2027

EviQ: Hierarchical Query-Conditioned Evidence Selection for Efficient VideoLLMs

Abstract

Video large language models (VideoLLMs) have achieved strong performance in video understanding, yet their large visual token budgets introduce substantial inference latency and memory consumption. Most existing training-free token compression methods rely primarily on visual cues, potentially overlooking query-critical evidence scattered across sparsely sampled frames, while some query-aware approaches incur additional computation through within-LLM attention or auxiliary models. Instead, we reveal that cosine similarity between projected visual tokens and embedded query tokens provides a lightweight relevance signal both across and within sparsely sampled frames. Building on this observation, we propose **EviQ**, a training-free pre-LLM token pruning framework that constructs a compact query-conditioned evidence set in two stages. **EviQ** first routes the global token budget across frames according to frame-level query relevance, and then selects tokens within each frame by jointly preserving informative visual context and promoting complementary coverage of query semantics. Extensive experiments on diverse VideoLLMs and multiple benchmarks demonstrate that **EviQ** achieves state-of-the-art accuracy-efficiency trade-offs among training-free token compression methods. Notably, **EviQ** preserves **99.9%** of vanilla LLaVA-OV-7B accuracy using only 25% of visual tokens, while achieving a 1.59 speedup and reducing GPU memory consumption by 8.7%. Code is released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.