acceptodds
Under review as a conference paper at ICLR 2027

Before You Ask: Learning to Proactively Initiate Interaction in Streaming Video

Abstract

Streaming video assistants require model to provide timely, context-aware support while remaining silent when intervention is unnecessary. Existing methods remain fundamentally reactive, relying on explicit user queries to trigger responses, thus failing to autonomously initiate assistance in dynamic real-world scenarios. To bridge this gap, we propose a zero-trigger paradigm that shifts the objective from merely deciding when to respond to autonomously determining whether, when, and how to intervene. Specifically, we replace the conventional binary speak/silence actions with four functional tokens: inquiry, response, notify, and silence. To realize this, we introduce SAIWen, a streaming multimodal model post-trained via a two-stage framework on SAIWen-Interact, a dataset of 26,549 training and 2,000 test videos with temporally aligned dialogues. During Supervised Fine-Tuning (SFT), our a Hierarchical Class-Balanced Loss decouples macro-level intervention triggering from micro-level functional routing, mitigating extreme functional token imbalance without degenerating into trivial always-silent policies. Subsequently, Reinforcement Learning (RL) with a Curriculum Advantage Annealing directly optimizes temporal alignment against our MII-Score, which rigorously evaluates proactive inquiry quality and notification precision while penalizing passive abstention. Experiments demonstrate that SAIWen significantly advances autonomous video assistance, improving intervention precision by 30.32% and notification timeliness by 14.90%, while strictly minimizing interruption costs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.