acceptodds
Under review as a conference paper at ICLR 2027

AudioLift: Audio-Driven Single Image to Full-Head Gaussian Avatar via Multi-view Video Generation

Abstract

Existing 3D talking avatar methods mainly rely on parametric face models including 3DMM or FLAME, whose compact motion representations limit the models ability to synthesize extreme expressions and facial accessories modeling. We introduce AudioLift, a two-stage framework that uses generated multi-view videos as the intermediate representation for audio-driven avatar generation without parametric face models or per-instance optimization. Our framework combines a synchronized multi-view video generation stage with an inconsistency-robust Gaussian reconstruction stage to recover coherent full-head geometry from a single portrait image. Extensive experiments demonstrate that AudioLift produces 3D talking avatars capable of capturing more extreme facial expressions than existing methods, with improved geometric completeness and strong multi-view consistency across the full viewing range.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.