acceptodds
Under review as a conference paper at ICLR 2027

Proact-VL2: Reinforcement Learning for Open-Ended Proactive Real-Time Streaming Interaction

Abstract

As AI companions increasingly participate in shared interactive experiences, such as co-watching and commenting on a game video with a user, they must continuously process streaming inputs to determine when and how to contribute. However, existing proactive real-time streaming models primarily optimize for predefined targets via supervised fine-tuning, such as flagging specific events or answering set queries, which falls short in open-ended scenarios, as it fails to ensure that the interaction trajectory remains contextually grounded, dynamically adapts to shifting conversational dynamics, and stays well-coordinated with others. We address this gap with reinforcement learning using a reward that assesses contextual grounding, adaptability, and human-AI coordination. Experiments demonstrate that this approach outperforms supervised learning and reference-based rewards, substantially improving proactive interaction quality and reducing multimodal grounding errors while preserving general video understanding capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.