acceptodds
Under review as a conference paper at ICLR 2027

Self-Focus: Native Active Perception for Streaming Interaction Models

Abstract

Deciding when to respond depends on the evidence a model has acquired. We introduce Self-Focus, a streaming interaction model that learns active perception natively within its interaction policy; thereby jointly deciding what evidence to acquire and when it is sufficient to respond. As the video stream progresses, Self-Focus constructs its own working context through active sensing, continuously, before and after a query arrives. We initialize the policy through supervised fine-tuning on our streaming episodes, then optimize it with GRPO over the policy's self-generated context. We introduce a delayed sensing credit to estimate each sensing action's advantage in subsequent interaction return. Self-Focus achieves state-of-the-art performance on both proactive and retrospective tasks across four benchmarks. The resulting query-agnostic context also improves offline long-video reasoning and provides a stronger initialization for iterative video agents, enabling better performance at fewer search iterations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.