StreamMind: Proactive Video Agents for Evolving User Intent
Abstract
Vision-language models are shifting from offline passive understanding toward online proactive interaction. However, prevailing methods largely rely on explicit requests while miss the implicit user intent (e.g., “where should I click next?” while a user navigate a website), meanwhile fail to update the state simultaneously as user requests might update continuously. These limitations make them less suited to intent-aware interaction in practical. To address these challenges, we introduce **StreamMind-Bench**, a new benchmark serving as testbed for user intent interaction, with 4.1K evaluation samples and 80K training samples spanning explicit, implicit, and dynamic intents. Building on this benchmark, we develop **StreamMind-VL**, a streaming proactive model that infers implicit user intents from continuous observations and dynamically tailors its interactions as these intents evolve. We further introduce **StreamMind-Harness**, which maintains user intents as persistent states with a fleixble lifecycles via a hierarchical memory to track the latest user intent and environmental state. Our experiments reveal three key findings: *(i)* StreamMind-VL preserves strong general video understanding while improving response timing and answer quality by 10.7% over other baselines; *(ii)* existing proactive models remain limited in understanding user intent and deciding when to respond, especially under implicit and dynamic intents; and *(iii)* StreamMind-Harness further improves response consistency and adaptation over sustained interaction, resulting in a 12.3% gain in proactive interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.