acceptodds
Under review as a conference paper at ICLR 2027

SetGain: Set-Conditioned Adaptive Evidence Selection for Long Video Understanding

Abstract

Long-form video question answering requires selecting a compact yet sufficient evidence set from thousands of frames. Existing methods follow predefined search procedures or optimize selection through trajectory-level outcomes, without modeling how much new information a candidate contributes beyond the evidence already collected. We propose SetGain, a set-conditioned evidence-selection framework that jointly learns which frame to select next and when to stop. A gain model conditioned on the candidate pool and the evolving selected set estimates each remaining frame's marginal contribution, trained by supervised initialization with marginal temporal-coverage targets followed by reinforcement learning with answer-quality and evidence-coverage rewards. On Haystack-LVBench, SetGain improves temporal F1 by 13.1% relative over a pointwise top- selector using the same candidate pool and evidence scores, with 16% fewer frames; end to end it reaches the temporal F1 and 8.2 points higher QA accuracy (15.7% relative) than the strongest comparable prior method, TimeSearch-R. It further improves average QA accuracy by 2.9 points (4.7% relative) over TimeSearch-R on VideoMME, MLVU, and LongVideoBench, and reduces end-to-end latency by 53.5% ( speedup) relative to T*, the fastest prior search method. These results demonstrate that explicitly modeling marginal evidence gain improves temporal localization, downstream answering, and inference efficiency. Code is available at https://anonymous.4open.science/r/SetGain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.