ConfabMuse: Structured Interaction Control for Multi-Person Talking Avatars
Abstract
Generating realistic multi-avatar conversational videos requires more than synchronizing multiple audio streams with visible avatars. A multi-avatar conversational generation framework must determine when each stream should influence motion, who should be animated by each stream, and how avatar behavior should evolve as the interaction unfolds. Existing methods largely treat multi-avatar generation as audio routing. They associate audio streams with avatars, but provide limited modeling of cross-stream turn dynamics and little structured control over how actions and relations evolve over time. To address this, we introduce ConfabMuse, a diffusion Transformer framework for controllable two-person conversational video generation from a reference image, two driving audio streams, and optional text guidance. ConfabMuse introduces a Turn-Context Audio Adapter for within and cross-stream speech modeling, Audio-Vision Binding for learned stream-to-avatar correspondence, and a Scene-Grounded Interaction Planner for structured control of evolving avatar actions and relations. Experiments across single and multi-person talking-avatar benchmarks show that ConfabMuse achieves state-of-the-art conversational interactivity, improves structured interaction control, and maintains strong audio-visual synchronization and generation quality. These results demonstrate the benefit of jointly modeling conversational timing, avatar correspondence, and interaction dynamics for controllable multi-person talking-avatar generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.