acceptodds
Under review as a conference paper at ICLR 2027

WATCH: Budgeted Active Perception and Visual Memory Control for Efficient Multimodal Reasoning

Abstract

Multimodal efficiency is often treated as a compression problem: process a fixed visual input, then reduce its context. As visual inputs grow in resolution and duration, this view overlooks a fundamental asymmetry: evidence must be acquired before it can be compressed, yet downstream compression cannot recover evidence that was never acquired. Active perception adapts what to observe; token reduction adapts what to retain. Separating these stages misses their causal dependence: the value of new evidence depends on what can survive bounded memory, while memory determines whether further perception is worthwhile. We introduce WATCH, a budget-conditioned closed-loop framework for sequential evidence allocation over bounded visual memory. At each step, it halts or acquires a view; only after new evidence arrives does it update bounded persistent memory. WATCH separates cumulative observation cost from memory capacity, enforces both constraints by construction, and learns the structured policy from downstream utility with reinforcement learning; the process admits a budgeted optimal-stopping interpretation and requires a single terminal call to the frozen language backbone. Across image and video benchmarks, WATCH improves the quality–efficiency frontier over passive, compression-only, and active-perception baselines while reducing both visual acquisition and terminal context; gains persist across budgets and backbones and translate into lower latency and memory. The anonymized code is available at https://anonymous.4open.science/r/WATCH/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.