ReDyFlow: Temporal-Pyramid Residual Flow Matching for Dyadic Conversational Avatar Generation
Abstract
Dyadic conversational avatar generation requires speech-aligned articulation and coordinated behaviors that are not uniquely determined by audio. We introduce ReDyFlow, a two-stage framework that jointly generates the speaking and listening head-motion sequences of both participants, without requiring observed partner motion as input. Given paired audio, speaking activity, and reference motion styles, a role-aware deterministic predictor constructs paired motion anchors through multi-scale temporal modeling and interaction between preliminary predictions. An anchor-conditioned joint flow then stochastically refines both trajectories with a temporal-pyramid velocity network that captures multi-resolution context and bidirectional participant interactions. We further introduce rollout-based dynamics supervision to regularize the temporal statistics of integrated samples. Experiments on the UniLS benchmark show that ReDyFlow achieves better results on the majority of quantitative metrics, while qualitative comparisons and user preferences indicate more natural speaking motion and listening reactions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.