acceptodds
Under review as a conference paper at ICLR 2027

REASON FIRST, RETRIEVE WHAT MATTERS: HYPOTHESIS-GUIDED KEYFRAME SELECTION FOR LONG VIDEO UNDERSTANDING

Abstract

Long video understanding requires models to recover and reason over sparse, temporally distributed evidence under limited frame budgets. Existing keyframe selection (KFS) methods typically retrieve query-relevant frames from a densely sampled candidate set before inference. Despite recent progress, existing KFS methods still have two key limitations. Global image–text matching introduces unnecessary computation when a small set of uniformly sampled frames already provides sufficient evidence. Moreover, selecting frames mainly by query relevance may miss evidence that distinguishes competing answers, and one-shot frame selection provides no opportunity to correct selection errors. We propose InferKFS, a training-free and plug-and-play framework that couples keyframe selection with iterative MLLM inference. InferKFS first reasons over sparsely sampled frames to construct an initial hypothesis space. The resulting answer distribution determines whether further frame selection is required; when activated, the retained hypotheses guide discriminative frame selection, followed by renewed inference and hypothesis refinement. This iterative process focuses additional visual computation on unresolved hypotheses while avoiding unnecessary global frame matching. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate consistent improvements across different MLLM backbones. InferKFS improves accuracy by up to 6.7 percentage points and achieves 72.9% accuracy on Video-MME with only 1.3% relative frame–hypothesis matching cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.