acceptodds
Under review as a conference paper at ICLR 2027

PIFBench: Evaluating and Improving Proactive Interaction in Streaming Video Understanding

Abstract

As AI systems move toward real-time human-AI interaction, evaluating models' ability to understand continuously arriving visual and audio information becomes increasingly important. In this context, streaming video understanding requires models to process incoming content incrementally and decide whether and when to respond. The ability to make this decision without an explicit user query is known as proactive interaction. Existing benchmarks for proactive interaction often omit audio or rely on laborious open-ended evaluation, and effective methods for improving this capability in omni-modal models remain underexplored. To address these gaps, we study proactive interaction from a full-stack perspective, covering evaluation, data construction, and model adaptation. For evaluation, we formalize Proactive Instruction Following (PIF), a task defined by providing an instruction before a video begins and requiring the model to process the video incrementally, remain silent when the specified condition is unmet, and produce the required response when it is met. Based on this task, we introduce PIFBench, a comprehensive benchmark for evaluating proactive interaction in streaming video understanding. PIFBench contains 1,993 items, covers both visual and audio information, and supports stable rule-based verification. Our evaluation shows that mainstream models have limited proactive interaction capability. Thus, we build Proactive-Instruct-8K, a structured training dataset for omni-modal models. For model adaptation, we make reusable modifications to the training framework and fine-tune Qwen3-Omni-30B-A3B-Instruct on Proactive-Instruct-8K, yielding Qwen3-Omni-Proactive, which improves accuracy on PIFBench from 13.79 to 19.42 under the same online evaluation setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.