acceptodds
Under review as a conference paper at ICLR 2027

CoFi: Compress First, Focus Later for Long Video Understanding

Abstract

Long-video understanding requires allocating a limited visual input budget to observations that are both temporally informative and relevant to the question. Existing training-free methods typically address this bottleneck through query-aware frame selection, visual compression, or adaptive combinations of the two. However, adaptive representation methods generally determine where to allocate visual capacity before explicitly accounting for what information a compressed view of the video is estimated to preserve. We introduce CoFi (Compress First, Focus Later), a training-free visual input construction framework that makes compression and focus acquisition sequentially dependent. CoFi first constructs a query-independent compressed temporal context and conditions a state–trajectory relational prior on this context to estimate the residual visual–temporal structure. It then introduces the question and sequentially acquires standalone focus observations that reduce query-relevant residual uncertainty, updating the residual structure after each acquisition. Under a fixed visual-unit budget, CoFi achieves the best performance on MLVU and LongVideoBench while remaining competitive with the best-performing frame-selection baseline on Video-MME. Ablation studies further show that compressed context and context-conditioned sequential focus acquisition play complementary roles, supporting the effectiveness of the proposed compress-first, focus-later design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.