acceptodds
Under review as a conference paper at ICLR 2027

Beyond Turn-Taking: Multimodal Continuous Full-Duplex Emotional Interaction

Abstract

Most existing multimodal conversational agents follow a turn-taking paradigm and typically remain silent while the user is speaking. This limits their ability to maintain perception while continuously generating multimodal feedback, such as facial expressions, conditioned on the user's evolving state and interaction context. We formulate this setting as multimodal continuous full-duplex emotional interaction, in which perception and response generation proceed concurrently. To support this task, we introduce **ContinuEI**, a dataset containing 30.24 hours of temporally aligned interactions from 1,072 dialogues and 12,303 sentences. It pairs users' visual-acoustic affective cues with system-side text, speech, and 3D facial motion, preserving the temporal correspondence between evolving user states and multimodal system responses. We further develop **FuSE**, an end-to-end **Fu**ll-duplex **S**treaming **E**motional interaction framework built upon Qwen3-Omni to provide a low-latency baseline for this task. FuSE incorporates an Actor module for continuous facial behavior generation, together with spatio-temporal attention and dual-stream conditional modeling. Experiments show that FuSE improves interaction naturalness and reduces response latency compared with conventional cascaded turn-taking baselines. These results support multimodal continuous emotional interaction as a promising task for studying continuous perception and multimodal response generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.