Aero Realtime: Duplex Video Interaction within a Unified Causal Stream
Abstract
The world continues to change while a video-language model generates a response. Continuous interaction requires new observations to enter during generation, yet turn-formatted execution defers them to a later input turn, while separate perception and response paths introduce additional coordination requirements. Can a video-language model support continuous input and output while preserving the training and incremental execution primitives of causal language models? We introduce Aero Realtime, a 4B model that places video, audio, silence, and lexical output on one causal timeline. At each approximately 80-ms audio slot, the model incorporates newly available visual states and predicts one lexical token or silence. Our contributions are threefold. 1) An append-only causal-stream architecture removes the native turn boundary, enabling continuous observation and generation within a single causal stream while preserving causal-language-model training and incremental execution. 2) We assemble 1.5M QA samples over 713K unique videos and address dense silence supervision with silence-label masking to adapt a pretrained VLM to the interface. 3) We provide parallel training and incremental serving as a distributed training and deployment implementation. Aero Realtime achieves 48.31 on OVOBench and records 84-ms median and 173-ms P95 processing lag over the first 20 minutes of a continuous stream on four NVIDIA A6000 GPUs. These results establish the learning and execution feasibility of duplex video interaction within a unified causal stream, with longer-run latency and autonomous response timing remaining limitations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.