Active Visual Perception and Reasoning for Long Video Understanding
Abstract
Long-video understanding requires identifying sparse, query-relevant evidence from a large amount of visually redundant content, making exhaustive video processing both inefficient and unnecessary. We present QSTAR, a Query-dependent Spatial-Temporal Active Reasoning framework for efficient long-video understanding without retraining. QSTAR casts video understanding as an iterative process of active perception, in which the model progressively concentrates its limited visual budget on moments that are most informative for answering the query. Starting from sparsely sampled frames, QSTAR exploits the VLM's internal visual reasoning signals to estimate query-dependent frame saliency and propagates these observations over the video timeline to refine a temporal sampling distribution. Successive rounds adaptively update the sampling distribution until the change between consecutive distributions falls below a threshold, after which a small set of informative frames is used for final multimodal reasoning. We further explore complementary sampling strategies and diversity-aware model fusion to improve robustness when different perceptual cues provide complementary evidence. We provide finite-round guarantees on reconstruction and induced sampling-distribution error. Extensive experiments on five video-understanding benchmarks, against 17 representative open- and closed-source VLMs and 10 methods specialized for long-video understanding, show that QSTAR consistently improves reasoning accuracy at substantially lower end-to-end latency. These results demonstrate that actively deciding where to look, conditioned on the query and the model's evolving reasoning state, provides an effective alternative to exhaustive long-video processing. Code is available at https://anonymous.4open.science/r/QSTAR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.