acceptodds
Under review as a conference paper at ICLR 2027

AnyTalker: Data-Efficient Scaling of Multi-Person Talking Video Generation

Abstract

Generating multi-person talking videos requires modeling not only per-speaker lip synchronization but also cross-speaker interactions, including gaze shifts, nods, and listening reactions. Existing methods often rely on hundreds or thousands of hours of curated multi-person video, which is costly to annotate and concentrated in domains such as studio podcasts. In this paper, we introduce AnyTalker, a data-efficient framework that shows how to build multi-person generation largely from abundant, open-domain single-person data. To this end, we concatenate independent single-person clips to teach the model to bind each audio stream to its corresponding face. To extend this learned binding beyond the two-person training setting, each DiT block uses an Audio-Face Cross Attention module with parameters shared across identity-audio pairs, allowing the number of identities to vary at inference time. We then refine listener behavior using 12 hours of real two-person conversations. Our evaluation combines human and MLLM assessments of listener reactions with emotion metrics and a calibrated responsiveness metric that compares listener motion with real videos. Despite training on at most two people, AnyTalker generates three- and four-person conversations with listener interactions. Across the evaluated benchmarks, it matches or surpasses methods trained on more real multi-person data in lip synchronization, visual quality, and human judgments of naturalness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.