acceptodds
Under review as a conference paper at ICLR 2027

Alive-Avatar: Streaming Audio-Video Avatar Generation with Persistent Identity

Abstract

Audio-video avatar generation is moving toward real-time, interactive streaming, where the core challenge is to sustain fast response, interactive control, and a persistent identity together over long sessions. We present ALIVE-AVATAR, a streaming audio-video avatar generation model that meets these demands. We first adapt the pretrained ALIVE base model through prompt-to-shot teacher forcing to enable promptable shot streaming for interactive control. Then, we design a distillation curriculum, initialized with consistency distillation and refined by our PCM-regularized DMD2, that attains high-quality generation in four sampling steps. At inference, our streaming pipeline uses Context Multiplicity Modulation to regulate attention to recent generated context, mitigating identity drift while retaining bounded historical memory. Together, these designs yield a four-step, 4-NFE model that, on our long-sequence benchmark, outperforms strong avatar baselines in identity consistency, face quality, and speech correctness, renders 480p audio-video at 24 fps faster than playback, and preserves identity even in one-hour rollouts. Building on ALIVE-AVATAR, we further construct a live-streaming agent that delivers planned content and answers viewer comments, demonstrating ALIVE-AVATAR's practicality for real-world avatar live streaming.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.