WATCH: Budgeted Active Perception and Visual Memory Control for Efficient Multimodal Reasoning
Abstract
Multimodal efficiency is often treated as a compression problem: process a fixed visual input, then reduce its context. As visual inputs grow in resolution and duration, this view overlooks a fundamental asymmetry: evidence must be acquired before it can be compressed, yet downstream compression cannot recover evidence that was never acquired. Active perception adapts what to observe; token reduction adapts what to retain. Separating these stages misses their causal dependence: the value of new evidence depends on what can survive bounded memory, while memory determines whether further perception is worthwhile. We introduce WATCH, a budget-conditioned closed-loop framework for sequential evidence allocation over bounded visual memory. At each step, it halts or acquires a view; only after new evidence arrives does it update bounded persistent memory. WATCH separates cumulative observation cost from memory capacity, enforces both constraints by construction, and learns the structured policy from downstream utility with reinforcement learning; the process admits a budgeted optimal-stopping interpretation and requires a single terminal call to the frozen language backbone. Across image and video benchmarks, WATCH improves the quality–efficiency frontier over passive, compression-only, and active-perception baselines while reducing both visual acquisition and terminal context; gains persist across budgets and backbones and translate into lower latency and memory. The anonymized code is available at https://anonymous.4open.science/r/WATCH/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.