Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Abstract
Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Finetuning a pretrained bidirectional model for this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context introduces exposure bias, causing errors and drift to accumulate over long rollouts. Second, conditioning on a global speech utterance cannot tell a causal generator which portion should be spoken next when only limited local context is available. We present Vorch-Streamer, a post-training framework that addresses these dilemmas and enables real-time long-form Text-to-Audio-Video (T2AV) streaming generation. We address the first dilemma through long-horizon Self Forcing, which exposes the model to its own long-term rollout distribution under the supervision of a bidirectional teacher. For the second dilemma, we introduce explicit speech progression control, where an external language model predicts discrete speech-planning tokens that condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the real-time efficiency while maintaining competitive audio-lip synchronization and strong identity preservation over long-form generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.