acceptodds
Under review as a conference paper at ICLR 2027

ConfabMuse: Structured Interaction Control for Multi-Person Talking Avatars

Abstract

Generating realistic multi-avatar conversational videos requires more than synchronizing multiple audio streams with visible avatars. A multi-avatar conversational generation framework must determine when each stream should influence motion, who should be animated by each stream, and how avatar behavior should evolve as the interaction unfolds. Existing methods largely treat multi-avatar generation as audio routing. They associate audio streams with avatars, but provide limited modeling of cross-stream turn dynamics and little structured control over how actions and relations evolve over time. To address this, we introduce ConfabMuse, a diffusion Transformer framework for controllable two-person conversational video generation from a reference image, two driving audio streams, and optional text guidance. ConfabMuse introduces a Turn-Context Audio Adapter for within and cross-stream speech modeling, Audio-Vision Binding for learned stream-to-avatar correspondence, and a Scene-Grounded Interaction Planner for structured control of evolving avatar actions and relations. Experiments across single and multi-person talking-avatar benchmarks show that ConfabMuse achieves state-of-the-art conversational interactivity, improves structured interaction control, and maintains strong audio-visual synchronization and generation quality. These results demonstrate the benefit of jointly modeling conversational timing, avatar correspondence, and interaction dynamics for controllable multi-person talking-avatar generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.