Self-Focus: Native Active Perception for Streaming Interaction Models
Abstract
Deciding when to respond depends on the evidence a model has acquired. We introduce Self-Focus, a streaming interaction model that learns active perception natively within its interaction policy; thereby jointly deciding what evidence to acquire and when it is sufficient to respond. As the video stream progresses, Self-Focus constructs its own working context through active sensing, continuously, before and after a query arrives. We initialize the policy through supervised fine-tuning on our streaming episodes, then optimize it with GRPO over the policy's self-generated context. We introduce a delayed sensing credit to estimate each sensing action's advantage in subsequent interaction return. Self-Focus achieves state-of-the-art performance on both proactive and retrospective tasks across four benchmarks. The resulting query-agnostic context also improves offline long-video reasoning and provides a stronger initialization for iterative video agents, enabling better performance at fewer search iterations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.