FlowTalker: Causal Streaming Audio-Driven 3D Head Animation via Factorized Motion Modeling with Rectified Flow
Abstract
Causally generating expressive 3D head animation as audio streams in is essential for real-time digital human interaction. However, most existing methods rely heavily on future audio context, leading to noticeable performance degradation when only limited or no look-ahead is available in real-time streaming inference. Current streaming-oriented approaches achieve low latency, yet their animation quality remains suboptimal, particularly in terms of expressive facial motion and long-term temporal consistency. We present FlowTalker, a framework for audio-driven 3D head animation under causal streaming inference. Pose, expression, and mouth motion evolve at different temporal scales and respond to distinct audio cues, but prior methods typically model them in a shared motion space. FlowTalker preserves these distinct motion dynamics by encoding them into separate token streams using spatially factorized vector quantization. We further introduce a Draft-then-Refine paradigm, in which a kinematics-aware autoregressive model predicts factorized motion tokens, and a selective rectified-flow module subsequently refines the expression and mouth latents to enhance expressive facial dynamics. Extensive experiments show that FlowTalker achieves state-of-the-art lip synchronization and facial motion quality under causal streaming inference, while showing the smallest performance degradation when transitioning from non-causal to causal settings. FlowTalker further enables high-fidelity photorealistic animation of 4D head avatars reconstructed from a single image.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.