From Scan to Focus: Learning a Value-Guided Adaptive Browsing Policy for Efficient Long-Video Reasoning
Abstract
Long-video reasoning requires Multimodal Large Language Models (MLLMs) to locate sparse evidence over extended temporal contexts while preserving the fine-grained visual details needed for reliable answers. Uniform sparse sampling can miss brief but decisive events, whereas dense processing incurs substantial visual-token and memory costs. To address this challenge, we introduce MindSight, a human-inspired adaptive browsing agent that treats long-video understanding as active perception. MindSight follows a scan–focus–reflect process: it first builds a low-cost global view to identify candidate intervals; conditioned on a candidate interval and the evidence acquired so far, it then compares STOP, LOW, and HIGH by their expected downstream returns, deciding whether to stop or inspect the interval at low or high visual fidelity by trading off new evidence against its observation cost. We learn this observation-value decision through paired Monte Carlo continuations at matched timestamps, and optimize multi-turn browsing trajectories with turn-balanced geometric-mean policy optimization (TB-GMPO). Experiments with two backbones on four long-video benchmarks improve accuracy by 1.2–12.0 percentage points (pp) over the corresponding base models, while reducing visual-token consumption and improving inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.