MTalk: Audio-Driven Multi-Person Multi-Shot Conversational Video Generation via Speaker-Identity Conditioning
Abstract
Audio-driven human video generation has progressed from animating single person to synthesizing multi-person conversations. Prevailing multi-person talking video frameworks rely heavily on spatially grounded audio binding, where speech features are hard-routed to target visual regions via spatial masks. However, this rigid spatial coupling fundamentally limits natural human interactions, frequently causing binding failures under dynamic movements, suppressing non-verbal listener dynamics, and failing during off-screen speech or shot transitions. To overcome these boundaries, we propose MTalk, a framework for multi-person multi-shot conversational video generation built on speaker-identity conditioning. Our key insight is to decouple who is speaking and when from visual localization. Specifically, Speaker-Aware Audio Modulation (SAM) encodes a supplied Speaker Activity Timeline (SAT) into time-aligned Speaker-Aware Embeddings (SAEs) that modulate acoustic features without per-person spatial masks. Temporal Context Propagation (TCP) preserves visual coherence across consecutive generation chunks, while Multi-Reference Visual Conditioning (MVC) supplies shot reference images and ID board via dedicated condition slots to support multi-shot transitions. Trained on diverse natural conversational videos, MTalk effectively handles complex turn-taking, off-screen speech, and multi-shot, achieving state-of-the-art performance on audio-driven video generation benchmarks. Project page: https://anonymous.4open.science/w/M2Talk-F011.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.