acceptodds
Under review as a conference paper at ICLR 2027

RI-Avatar: Structure-Aware Real-Time Streaming for Personalized Avatars

Abstract

Speech-driven avatars require coherent real-time generation under a tight latency budget while maintaining temporal continuity across streaming chunks. Existing few-step diffusion avatars still spend uniform denoising computation over regions whose motion importance is highly non-uniform, while imperfect state propagation can introduce visible artifacts at chunk boundaries. We present RI-Avatar, a real-time streaming framework centered on two complementary ideas. First, structured region-aware denoising (SRAD) anchors each chunk with a full-frame step and concentrates subsequent refinement on speech-sensitive facial regions. Second, a codec-inspired Pixel State re-encodes the decoded RGB tail of each chunk as causal state for the next, preserving identity, pose and expression across chunk boundaries. On the unified HDTF/VFHQ evaluation, the SRAD route achieves a 1.45× resident speedup over the official native FlashHead Lite baseline while providing a substantially better quality–efficiency trade-off than matched token-merging and pruning controls. With the final RGB Pixel State runtime optimizations enabled, a separate representative 22.4-s continuous-speech stream reduces resident generation time from 8.149 s to 4.370 s, reaching 1.865× speedup on a single A800. As a deployment extension, we additionally implement a stateful identity handoff path with CPU-resident identity bundles, bounded active/standby GPU slots, and epoch-tagged media-timeline cutover; an A→B→A serving smoke verifies bundle and timestamp ordering during continuous generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.