acceptodds
Under review as a conference paper at ICLR 2027

DuplexTTS: A Scalable Framework for Full-Duplex Conversational Speech Synthesis and Evaluation

Abstract

The scarcity of full-duplex conversational speech data limits the development of natural spoken interaction models. Conventional text-to-speech systems primarily generate isolated utterances or alternating turns, leaving the joint generation of conversational content and interaction timing underexplored. Motivated by advances in native full-duplex speech generation, we introduce DuplexTTS, a scalable framework for synthesizing and evaluating full-duplex conversational speech. DuplexTTS supports two synthesis routes: native autonomous generation, where a single speech model generates two role-conditioned streams through shared acoustic feedback, and rule-based orchestration, where scripted utterances are synthesized and arranged into turn exchanges, backchannels, interruptions, and side talk. A unified pipeline connects synthesis with dialogue-level quality screening, additional source-utterance checks for rule synthesis, interaction analysis, and downstream evaluation. Using this framework, we construct over 1,800 hours of screened Chinese and English dialogue. Human listeners rate the generated dialogues favorably for naturalness and meaningfulness. Training Moshi on the two corpora yields different conversational behaviors. The mixed-data model achieves 16.60 percentage points higher macro expected-behavior accuracy on Full-Duplex-Bench than the original model. These results establish DuplexTTS as a practical framework for expanding full-duplex training data and studying how synthesis and curation shape spoken interaction. Audio demos are available at https://duplextts.github.io.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.