acceptodds
Under review as a conference paper at ICLR 2027

SteadyTalk: Remembering Every Speaker for Long-Horizon Audio-Video Generation

Abstract

Digital humans often appear in long videos with several speakers, and each speaker must keep the same voice and face throughout. Joint autoregressive generators drift in timbre and color over such long videos and forget a speaker’s voice and face after a long silence. In audio-driven generators, the given audio dominates the output and limits control over the scene. To address these issues, we propose SteadyTalk, a long audio-video generation framework for single- and multi-person scenes, built upon three modules. First, Audio-First Asynchronous Generation drives the joint model with its own audio. A frozen audio branch generates the audio of each chunk first, and the video branch follows both this audio and the text prompt. Second, Multi-Person Audio-Visual Memory stores key audio segments and frames for each participant and recalls them with the current prompt, so a returning speaker keeps the same voice and face. Third, Reference-Anchored Drift Correction aligns the color statistics of each chunk to the opening frames and its audio latent statistics to the speaker’s first utterance. On one-minute single- and two-person conversations, SteadyTalk outperforms existing methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.