DUET: Audio-Driven Human Video Synthesis with Voice-Face Dual Harmonization
Abstract
Existing audio-driven human video generation predominantly focuses on lip synchronization, yet often overlooks the timbre mismatch between the driving voice and the depicted speaker. We propose DUET, a voice-face dual-conditioned talking video generation model that animates the face with lips synchronized with the driving audio while adaptively updating the audio timbre to match the speaker appearances. We formulate this task within a joint audio-video generation framework: the driving utterance provides linguistic and temporal cues, while the conditional image determines the vocal identity of the generated speech. We incorporate a Whisper-based audio injection branch into a pretrained joint audio-video generative model and train it on timbre-perturbed audiovisual pairs to facilitate both lip synchronization and timbre reasoning through cross-modal learning. For multi-person conversational scenes, we introduce a spatio-temporal masking mechanism that associates each audio condition with its designated face region and speaking interval. On the speaker-identity-disjoint EMTD benchmark, our method obtains the best or tied-best scores on the evaluated voice-conversion metrics. On EMTD and InterActEyes, it maintains competitive visual quality and lip synchronization. In a perceptual comparison with a DreamVoice+MultiTalk cascade on EMTD, participants prefer our method in 64.29% of responses for timbre–identity consistency and 53.97% for overall audiovisual coherence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.