Beyond Turn-Taking: Multimodal Continuous Full-Duplex Emotional Interaction
Abstract
Most existing multimodal conversational agents follow a turn-taking paradigm and typically remain silent while the user is speaking. This limits their ability to maintain perception while continuously generating multimodal feedback, such as facial expressions, conditioned on the user's evolving state and interaction context. We formulate this setting as multimodal continuous full-duplex emotional interaction, in which perception and response generation proceed concurrently. To support this task, we introduce **ContinuEI**, a dataset containing 30.24 hours of temporally aligned interactions from 1,072 dialogues and 12,303 sentences. It pairs users' visual-acoustic affective cues with system-side text, speech, and 3D facial motion, preserving the temporal correspondence between evolving user states and multimodal system responses. We further develop **FuSE**, an end-to-end **Fu**ll-duplex **S**treaming **E**motional interaction framework built upon Qwen3-Omni to provide a low-latency baseline for this task. FuSE incorporates an Actor module for continuous facial behavior generation, together with spatio-temporal attention and dual-stream conditional modeling. Experiments show that FuSE improves interaction naturalness and reduces response latency compared with conventional cascaded turn-taking baselines. These results support multimodal continuous emotional interaction as a promising task for studying continuous perception and multimodal response generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.