acceptodds
Under review as a conference paper at ICLR 2027

DES: Discriminative Evidence Selection for Long-Video Question Answering

Abstract

Answering questions about long videos requires identifying relevant visual evidence distributed over time. With a limited frame budget, multiple-choice video question answering also requires evidence that distinguishes among candidate answers. However, question-relevant frames may provide little discriminative information, and several high-scoring frames may depict the same content. We introduce Discriminative Evidence Selection (DES), a training-free framework that estimates frame utility by combining semantic relevance, differences in frame–option matching, object and action cues, and local temporal evidence. DES constructs a candidate pool from local score maxima, temporal segment representatives, and globally high-scoring frames. It then selects a diverse subset using maximal marginal relevance, progressively relaxing the minimum temporal gap between selected frames as needed to fill the frame budget. The selected frames are provided in chronological order to a frozen video-language model. Experiments on four benchmarks with frame budgets of 8, 16, and 32 show that DES outperforms uniform sampling in most settings. With LLaVA-OneVision-7B and eight input frames, DES improves accuracy over BOLT by 2.0 and 3.3 percentage points on Video-MME and MLVU, respectively. Under the same eight-frame budget, evaluations on Video-MME without subtitles using five additional video-language models show accuracy gains of 3.2–6.1 percentage points over uniform sampling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.