acceptodds
Under review as a conference paper at ICLR 2027

Chorus: Extending Single-Speaker TTS to Multi-Speaker Conversation Generation

Abstract

Multi-speaker conversational speech generation synthesizes a conversation among several speakers from its transcript. Most existing systems generate the conversation as a single-channel mixture, which entangles the speakers. This makes the generated speech difficult to use for downstream tasks that require a separate channel for each speaker. It also complicates per-speaker evaluation, which requires diarization to attribute each speech segment to the corresponding speaker. Systems that generate one channel per speaker address these issues but are limited to two speakers and one form of timing control. We introduce Chorus, which assigns a weight-shared single-speaker text-to-speech (TTS) model to each speaker and connects the models through a communication block whose computation does not depend on the number of speakers. Chorus therefore supports an arbitrary number of speakers, chosen at inference time. Chorus retains the infilling objective of its base model, F5-TTS. Each model receives the audio of its channel, a mask over the frames to generate, and a text input in which special tokens mark the turns of that speaker. Within this formulation, four tasks are obtained by choosing which frames are masked and how the transcript is structured: generation of all channels from a transcript, generation of one channel conditioned on the recorded audio of the others, chunk generation of long conversations conditioned on the previous chunk, and generation with turn timing specified either by explicit start and end times or by the turn order alone. On the two-speaker ZipVoice-Dialog test set, Chorus is competitive with the baselines on intelligibility and speaker similarity, although it is fine-tuned on less conversational speech. From the turn order alone, it produces backchannels and interruptions whose timing a turn-taking judge scores above the ground truth. The same model generates two- to four-speaker conversations on AMI with a rise of one point in per-channel word error rate, making it, to our knowledge, the first system that generates one channel per speaker for more than two speakers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.