acceptodds
Under review as a conference paper at ICLR 2027

NavIAgent: Coarse-to-Fine Active Frame Selection for Long-Form Video Understanding

Abstract

Long videos contain vast temporal search spaces, yet answer-critical visual evidence may be confined to a few brief moments. Under limited frame budgets, effectively locating such sparse evidence requires both broad temporal coverage and fine-grained temporal precision. This makes it challenging to search the long temporal horizon efficiently while preserving the resolution needed to capture decisive evidence. To address this challenge, we propose NavIAgent, a coarse-to-fine active frame selection framework that combines global temporal localization with fine-grained evidence acquisition. NavIAgent operates through two complementary stages. (1) At the coarse stage, a lightweight multimodal event graph indexes structured segment memories and rapidly localizes query-relevant candidate intervals. (2) At the fine stage, we deploy Strategic Memory Informed Lens for Evidence (Smile), an active agent optimized via reinforcement learning with dual-scale advantage and Free Energy-inspired intrinsic rewards. By formulating keyframe selection as a Partially Observable Markov Decision Process (POMDP), Smile performs discrete temporal navigation within candidate intervals and progressively acquires answer-relevant evidence from intermediate observations. Theoretical analysis shows that active policies have a more expressive hypothesis space than passive sampling, while unconstrained discrete navigation suffers from asymptotic temporal resolution dilution as video length increases, providing a formal motivation for our coarse-to-fine design. Extensive experiments on Video-MME, LongVideoBench, and LVBench across multiple vision-language backbones demonstrate consistent accuracy improvements of up to 11.6 percentage points, while substantially reducing visual frame consumption and inference latency. Our code is in the supplementary materials.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.