acceptodds
Under review as a conference paper at ICLR 2027

DUET: Audio-Driven Human Video Synthesis with Voice-Face Dual Harmonization

Abstract

Existing audio-driven human video generation predominantly focuses on lip synchronization, yet often overlooks the timbre mismatch between the driving voice and the depicted speaker. We propose DUET, a voice-face dual-conditioned talking video generation model that animates the face with lips synchronized with the driving audio while adaptively updating the audio timbre to match the speaker appearances. We formulate this task within a joint audio-video generation framework: the driving utterance provides linguistic and temporal cues, while the conditional image determines the vocal identity of the generated speech. We incorporate a Whisper-based audio injection branch into a pretrained joint audio-video generative model and train it on timbre-perturbed audiovisual pairs to facilitate both lip synchronization and timbre reasoning through cross-modal learning. For multi-person conversational scenes, we introduce a spatio-temporal masking mechanism that associates each audio condition with its designated face region and speaking interval. On the speaker-identity-disjoint EMTD benchmark, our method obtains the best or tied-best scores on the evaluated voice-conversion metrics. On EMTD and InterActEyes, it maintains competitive visual quality and lip synchronization. In a perceptual comparison with a DreamVoice+MultiTalk cascade on EMTD, participants prefer our method in 64.29% of responses for timbre–identity consistency and 53.97% for overall audiovisual coherence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.