acceptodds
Under review as a conference paper at ICLR 2027

From Scan to Focus: Learning a Value-Guided Adaptive Browsing Policy for Efficient Long-Video Reasoning

Abstract

Long-video reasoning requires Multimodal Large Language Models (MLLMs) to locate sparse evidence over extended temporal contexts while preserving the fine-grained visual details needed for reliable answers. Uniform sparse sampling can miss brief but decisive events, whereas dense processing incurs substantial visual-token and memory costs. To address this challenge, we introduce MindSight, a human-inspired adaptive browsing agent that treats long-video understanding as active perception. MindSight follows a scan–focus–reflect process: it first builds a low-cost global view to identify candidate intervals; conditioned on a candidate interval and the evidence acquired so far, it then compares STOP, LOW, and HIGH by their expected downstream returns, deciding whether to stop or inspect the interval at low or high visual fidelity by trading off new evidence against its observation cost. We learn this observation-value decision through paired Monte Carlo continuations at matched timestamps, and optimize multi-turn browsing trajectories with turn-balanced geometric-mean policy optimization (TB-GMPO). Experiments with two backbones on four long-video benchmarks improve accuracy by 1.2–12.0 percentage points (pp) over the corresponding base models, while reducing visual-token consumption and improving inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.