acceptodds
Under review as a conference paper at ICLR 2027

From Relevance to Representatives: Equal-Rate Evidence Quantization for Long Video Understanding

Abstract

Long video understanding often requires selecting useful visual observations under a limited frame budget. Uniform sampling distributes observations by elapsed time, whereas global relevance ranking can concentrate them in locally redundant high-scoring regions. In this paper, we propose Equal-rate Evidence Quantization (E²Q), a training-free frame selector for long video understanding that separates where to allocate the budget from which observation to retain. E²Q uses normalized frame–question relevance to allocate the frame budget, partitioning the ordered candidates into consecutive intervals with approximately equal relevance mass and retaining the highest-scoring frame from each interval. A budget-dependent projection handles overly concentrated candidate weights before partitioning, enabling relevance-adaptive intervals without an additional relevance–coverage coefficient. We characterize this construction and bound frame-level representation error in terms of within-interval variation and interval-mass imbalance. Across four long-video question answering benchmarks and three models, E²Q achieves the highest accuracy among training-free selectors in ten of the twelve model–benchmark settings. Controlled shared-score comparisons further show consistent gains when the relevance signal is fixed and isolate the respective contributions of relevance-mass allocation and within-interval representative selection to these gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.