NextAvatar: Next-State Planning for Streaming Avatar Video Generation
Abstract
Streaming avatar generation synthesizes video chunk by chunk, conditioned on visual history and text instructions. Each chunk typically covers about one second, often too short to complete an action, so a single instruction spans multiple chunks during training and inference. Coherent action progression therefore requires the generator to account for what the avatar has already done. However, we observe that streaming generators often stagnate before executing an action or repeat it after completion. To investigate these failures, we analyze the attention mechanisms that incorporate visual history and text instructions into generation. We identify history-state confusion: attention to both conditions remains largely unchanged across distinct execution states under the same instruction, failing to adequately distinguish action progress and leading to stagnation and repetition. To address this problem, we propose NextAvatar. We use a frozen pretrained multimodal large language model to interpret the history state and active instruction, then provide an explicit next-state visual plan through state-conditioned attention. To reduce appearance deviations in the planned state, we introduce Reference-Grounded State Anchoring, which refines its visual features through asymmetric attention to fixed initial-reference features. We further introduce the Multi-Instruction Responsive Avatar Benchmark (MIRA-Bench), comprising 2,000 expression–action intervals to evaluate Human Fidelity, Instruction Responsiveness, Video Quality, and Action Progression. With a 1.3B generator backbone, NextAvatar achieves the strongest Instruction Responsiveness among the evaluated methods and reduces combined stagnation and repetition from the best external baseline's 52.5% to 30.5%, while maintaining competitive Human Fidelity and Video Quality. A demo page is available at https://anonymous.4open.science/w/nextavatar_site/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.