ALIVE: Making Streaming Video Models Interactive
Abstract
Streaming video LLMs offer a natural foundation for multimodal interactive AI that continuously perceives the physical world and interacts with users. Yet existing models remain largely passive, responding to individual requests in isolation. To bridge this gap, we study how to transform streaming video LLMs from passive observers into interactive assistants. We formulate interactive streaming video understanding around three core capabilities: proactive monitoring, multi-task coordination, and feedback seeking. Based on this formulation, we introduce ALIVE-234K, a large-scale dataset for post-training, together with FSI-Bench, a comprehensive benchmark for evaluating interactivity in streaming video LLMs. We further investigate post-training strategies and introduce ALIVE, a family of models that equips offline video LLMs with rich interactive behaviors without requiring specialized model architectures. Across multiple streaming video benchmarks, ALIVE achieves highly competitive performance, preserving strong video understanding while enabling rich interactions over evolving video streams. Our results demonstrate that interactive behaviors can emerge through appropriate data and post-training, providing a path toward multimodal assistants that continuously perceive and interact with users as events unfold.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.