acceptodds
Under review as a conference paper at ICLR 2027

Edge-Avatar: Pose-Controlled Real-Time Speech-to-Video on a Single Consumer GPU

Abstract

Audio-driven human animation reaches high fidelity, but its many-step guided sampling is far too slow for an interactive avatar. We present Edge-Avatar, a speech-to-video model for controllable audio-driven human animation that runs in real time at 25 FPS on a single consumer GPU under pose control. Starting from a bidirectional 2B diffusion transformer that draws text and audio conditioning from a single Qwen2.5-Omni model, two training stages convert it into a chunk-autoregressive generator, built on recent streaming works and extended with an area-adaptive face loss and memory noising. Pose control is trained rather than imposed at inference, through a pose track that may begin or end partway through a chunk, so a gesture is entered and left smoothly. An optional training-free reset returns the avatar to the reference image mid-stream and clears the carried memory. Deployment pairs the few-step sampler with FP8 linear layers, block-sparse attention, compiled execution, and a distilled decoder, so that denoising and decoding fit inside the playback interval of each chunk at SD resolution on a single NVIDIA RTX 5090. Two optional pieces are independent of each other, a cheaper opening denoising step carried in the frequency domain and a further-upscaling decoder, and together they deliver HD inside that same budget. On common speech-to-video benchmarks the result leads the open-source real-time avatar generation systems it is measured against, including ones an order of magnitude larger and ones that reach real time only on multi-GPU nodes, and a side-by-side study against the three closest of them comes out in its favour.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.