Bounding Latency Contribution of a Dyadic Avatar System
Abstract
Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. Existing work reports throughput or aggregate latency but does not show how generation and serving interact to shape the latencies a user actually experiences. We introduce XXXX-1, a compact autoregressive flow-matching motion model together with an end-to-end system for live dyadic avatar interaction. We show how to turn chunk-based generation into a real-time conversation between a user and an avatar. We analytically derive the system's contribution to two user-facing latencies and validate it with two commercial voice agents. Further experiments demonstrate that XXXX-1 achieves leading performance among the compared dyadic models across visual quality, lip synchronization, and listening behavior, while its inference runtime operates in real time on data-center and consumer GPUs. We release the https://anonymous.4open.science/r/bounding-dydactic-latency/anonymized code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.