Vision as Efficient Information Gathering
Abstract
Modern vision-language models perceive every pixel uniformly across the frame(s), spending compute on regions irrelevant to the task. While dynamic-resolution and adaptive-tokenization methods adapt the token budget to the input, they remain task-agnostic and must still process the full image at least once. We introduce ActiveLook, which casts visual perception as a decision-making problem. Starting from a coarse glance, we learn a policy that actively decides, turn by turn, where to look (bounding box, page index, temporal window) and how much to look (scale, frame rate), until it has gathered enough evidence to answer, optimizing accuracy per FLOP. The obstacle is supervision: ground-truth look sequences do not exist. We bootstrap this from a frozen teacher's rollouts, keeping trajectories that verifiably lead to the correct answer, and augment them with synthetic trajectories derived from ground-truth annotations. We distill these trajectories with KL-anchored supervised fine-tuning, followed by reinforcement learning with a compute reward that favors efficient looks across the action space: number of looks, bounding box, scale, temporal window, and frame rate. We gate the reward on correctness and schedule the gate to avoid sparse rewards early in training. ActiveLook learns a single policy spanning images, multi-page documents, and videos, that advances the accuracy-FLOPs Pareto frontier across all regimes: matching the base model's accuracy with less compute on images and long documents, and less on video. At matched input budgets, ActiveLook is points more accurate on average while using 27% less compute.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.