VESTA: Timbre-Faithful and Lip-Synced Joint Talking Audio-Video Generation
Abstract
Recent joint talking audio–video generation produces realistic human videos, yet falls short on both the voice and the lips: the timbre is uncontrolled rather than specified, and the lips stay imperfectly synchronized. Existing timbre-injection methods feed the reference into the token stream, which in a coupled generator leaks time-varying content and unsettles the lips; they supervise both modalities with reconstruction losses alone, and serve every speaker with one fixed pass. We present VESTA (Video-audio Expert-guided Synchronized Timbre-faithful Avatar), which injects timbre off the time axis and supervises perception directly on the decoded face and voice, through three components. (1) StatShift modulates the audio branch's channel statistics (AdaLN) uniformly across time with a reference speaker embedding; carrying no time-varying signal, it controls timbre while empirically preserving the content the lips track. (2) RewardSteer adds the perceptual supervision joint generators lack: a timbre reward, and for lip-sync a noise-aware critic scored on a face-localized decode of one-step predictions, since both streams are generated and only the face carries the signal. (3) TimbreTune is an optional per-sample test-time step for the hard-speaker tail. On the open-source EMTD benchmark, VESTA raises reference-timbre similarity by nearly 50% (from 0.40 to 0.60), a gain ablations attribute to the mechanism rather than the data; it also leads on lip-sync (Sync-C 7.43 vs. 6.55) and stays competitive on PQ, CU and FVD. TimbreTune adds +6.2%, most on the hardest speakers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.