DuplexAct: Benchmarking Natural and Controllable Interaction Behaviors in Real-Time Full-Duplex Speech Systems
Abstract
Real-time full-duplex speech systems can listen and speak simultaneously, but natural interaction also requires them to adapt their behavior to the conversational context and users’ spoken instructions. We introduce DuplexAct, a benchmark for evaluating how naturally and controllably these systems interact. DuplexAct covers six interaction families, each evaluated on two tracks: Capability and Command. The Capability track tests whether a system acts appropriately in the conversational context under a fixed policy for each family. The Command track evaluates whether a system follows users’ spoken instructions and adapts its behavior when those instructions are updated. We build 512 base cases in Chinese and English from real customer-service conversation records and validate the resulting dataset through human review. Our real-time evaluation protocol accommodates different system interfaces, assesses interaction behavior from actual audio playout traces, and reports correctness and latency separately. Most evaluated systems exceed 90% success in initiating responses to completed requests in both languages under the Capability track. In contrast, success rates reach at most 32.1% for waiting through a mid-utterance pause before responding, and 30.8% for timely intervention after factual errors. These behaviors remain challenging in the Command track. These findings highlight context-appropriate behavior and reliable spoken-policy execution as key challenges for natural and controllable full-duplex interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.