SteadyTalk: Remembering Every Speaker for Long-Horizon Audio-Video Generation
Abstract
Digital humans often appear in long videos with several speakers, and each speaker must keep the same voice and face throughout. Joint autoregressive generators drift in timbre and color over such long videos and forget a speaker’s voice and face after a long silence. In audio-driven generators, the given audio dominates the output and limits control over the scene. To address these issues, we propose SteadyTalk, a long audio-video generation framework for single- and multi-person scenes, built upon three modules. First, Audio-First Asynchronous Generation drives the joint model with its own audio. A frozen audio branch generates the audio of each chunk first, and the video branch follows both this audio and the text prompt. Second, Multi-Person Audio-Visual Memory stores key audio segments and frames for each participant and recalls them with the current prompt, so a returning speaker keeps the same voice and face. Third, Reference-Anchored Drift Correction aligns the color statistics of each chunk to the opening frames and its audio latent statistics to the speaker’s first utterance. On one-minute single- and two-person conversations, SteadyTalk outperforms existing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.