acceptodds
Under review as a conference paper at ICLR 2027

Segments Divide, Frames Decide: Composable Relevance for Long Videos from QA Alone

Abstract

A vision-language model (VLM) sees a long video only through a handful of frames, so the choice of frames determines what it can answer. Queries differ in the temporal scale of the evidence they need, yet existing relevance signals operate at a single, fixed scale. Selectors that require no temporal grounding annotations score each frame in isolation, missing motion and event structure. Supervised grounding methods capture this structure but require such annotations and expensive model passes. This paper introduces SPAN: Segment relevance via Propagated, Annotation-free, Nested composition, which scores relevance at multiple temporal scales without grounding labels, using a query-conditioned composition operator that propagates relevance from base clips to arbitrarily coarse segments. Its weighting and composition functions read only the defined base-clip representations and the query to construct segments and obtain their relevance, so the same operator applies unchanged across scales. Self-distillation learns this operator from video-question-answer tuples alone: composed clips must reconstruct the video's own representation. Finally, a nested, tree-structured knapsack allocates SPAN's frame budget across the segment hierarchy. Across four long-video benchmarks and two VLM families, SPAN improves on uniform sampling by an average of 5.47 points and beats the strongest annotation-free baseline in 18/20 settings. It recovers over 95.0% of the performance achieved by a selector that runs a VLM during selection, without running one itself. For long-video question answering, this makes query-aware frame selection a lightweight plug-in step, thereby letting any frozen VLM answer from relevant evidence without temporal grounding annotations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.