acceptodds
Under review as a conference paper at ICLR 2027

DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model

Abstract

Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. This high latency critically prevents the immediate, non-verbal feedback required for a realistic listener. To address this, we present DyStream, a flow matching-based autoregressive model that could generate video in real-time from both speaker and listener audio. Our method contains two key designs: (1) we adopt a stream-friendly autoregressive framework with flow-matching heads for probabilistic modeling, and (2) we minimize reliance on future context by using a causal encoder with a minimal (e.g., 60ms) lookahead module. We analyzed alternative causal strategies, including distillation and generative encoders. Our results show that a simple lookahead module is both simple and effective. Extensive experiments show that \ourmethod could generate video within 34ms per frame, guaranteeing the entire system latency remains under 100ms. Besides, it achieves state-of-the-art lip-sync quality, with offline and online LipSync Confidence scores of 8.13 and 7.67 on HDTF, respectively. The model, weights and codes are available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.