acceptodds
Under review as a conference paper at ICLR 2027

ERTalk: Autoregressive Speech-Driven 3D Face Animation Guided by Expressive Reference Motion

Abstract

Speech-driven 3D facial animation aims to generate natural and expressive facial motion with accurate lip synchronization. Discrete emotion labels alone provide limited information about fine-grained facial deformations. We introduce ERTalk, a reference-conditioned autoregressive model that uses speech and reference motion to generate synchronized lip movements and reference-guided expressions in a learned deformation space. Multi-State Reference Conditioning (MRC) augments the global reference prefix with layer-wise access to frame-level expression states, while differentiable decoding enables vertex and velocity supervision over all decoded vertices in valid frames. Our pretrained SMPL-X-based representation encodes facial expression separately from the target identity geometry. This representation allows the generated motion to be transferred across identities without rerunning the motion generator. Extensive evaluations demonstrate that our method outperforms existing approaches in both lip-sync accuracy and emotion control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.