acceptodds
Under review as a conference paper at ICLR 2027

UniVIA: A Unified Video Intelligence Architecture for Offline Understanding and Duplex Streaming Interaction

Abstract

Streaming video interaction requires a model to keep perceiving an evolving video while deciding when to respond and generating answers. Existing chunk-based formulations couple response timing with answer generation: fine chunks split each answer across observation updates, departing from the complete-answer supervision of offline video-language training, while coarse chunks sacrifice the temporal granularity of response decisions. We present UniVIA, a unified video intelligence architecture that serves offline video understanding and full-duplex streaming interaction with a single shared vision-language model. At inference, an asynchronous fork-and-commit mechanism lets a persistent parent stream keep perceiving the video while each response decision forks a child stream that generates the complete answer from a fixed context snapshot and commits it back as memory. For training, we split each streaming interaction into response-ended samples that reproduce the parent and child states of inference, and propose a Segment-Balanced Loss that supervises the silence and response decisions of each query with balanced class weights. UniVIA achieves the best results among open-source models on seven of ten tasks across StreamingBench, OVO-Bench, and OmniMMI, with the largest gains on proactive tasks, keeps the perception delay added by answering at 63–94 ms at 1 FPS regardless of answer length, and achieves the best results on nine of eleven offline video metrics among all evaluated models; with the Qwen3-VL-8B backbone, it improves over the base model on eight of these metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.