acceptodds
Under review as a conference paper at ICLR 2027

Continuo: A Large-Scale Dataset for Controllable Long-Form Speech Generation

Abstract

Generating long-form speech requires models to keep content, speaker identity, turn-taking, and expressive style consistent over time. Yet most open training corpora contain examples lasting only tens of seconds and open long-form benchmarks cover at most two speakers per recording. We introduce Continuo, a 229K-hour multilingual dataset for controllable long-form speech generation. From in-the-wild data, Continuo constructs short utterances, single-speaker long-form speech, and multi-speaker dialogues, preserving transcripts, speaker identities, pauses, and acoustic transitions. Examples longer than five minutes account for 50.0% of single-speaker long-form hours and 66.8% of dialogue hours. To support expressive control, Continuo provides natural language instructions for 4.77M examples and nonverbal-vocalization tags for 933K examples. We also introduce Continuo-testset, 23 hours of bilingual recordings of up to 30 minutes and four speakers, each speaker paired with reference audio. We train autoregressive (AR) and non-autoregressive (NAR) models on Continuo. Our 0.6B-backbone NAR model achieves the lowest Chinese CER and English WER on TTSD-Eval (5.27%/9.93%) among compared systems, while our AR model is on par with Qwen3-TTS-12Hz-VD on InstructTTSEval (zh/en: 76.2/78.3 vs. 77.1/77.9). These results validate Continuo as a comprehensive data foundation for long-form speech generation and minute-scale expressive control. Demo: https://continuo6.github.io/continuo-demo/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.