StreamPro: Timing-Aware Evaluation and Effective Training for Proactive Streaming Video Understanding
Abstract
Proactive streaming video understanding requires models to continuously monitor video streams and autonomously determine whether, when, and what to respond based on evolving user needs and visual context. Despite recent progress, both the evaluation and training of proactive models remain challenging. Existing benchmarks primarily focus on event-triggered responses, leaving anticipatory assistance underexplored, while their evaluation protocols often fail to jointly account for task-specific timing requirements, semantic correctness, and response frequency. Training also remains challenging: SFT suffers from severe imbalance between silence and response signals, whereas RL faces sparse response-level rewards and the challenge of optimizing interdependent response sequences. To address these challenges, we introduce StreamPro, a unified benchmark-and-training framework for proactive streaming video understanding. We first develop StreamPro-Bench, which evaluates three complementary dimensions of proactive capability: Perceptual Understanding, Temporal Reasoning, and Anticipatory Assistance. Its StreamPro-F1 metric jointly evaluates semantic correctness and response timing, while accounting for both excessive and missed responses. We further propose a two-stage training framework that combines class-balanced supervised fine-tuning with reinforcement learning. Specifically, CB-Stream Loss alleviates supervision imbalance during SFT, while GRPO with a refined Response-Matching (RM) reward and a rubric-based Sequence-Quality (SQ) reward provides denser feedback and encourages coherent response sequences. We construct StreamPro-SFT-63K and StreamPro-RL-3K to support the two training stages. Experiments show that StreamPro substantially improves proactive streaming video understanding: it achieves an average StreamPro-F1 of 32.8 on StreamPro-Bench, compared with 9.2 for the strongest existing baseline, while maintaining strong performance on real-time streaming video understanding, reaching 79.8 on StreamingBench-RTVU.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.