acceptodds
Under review as a conference paper at ICLR 2027

Continuous Spoken Dialogue Language Models

Abstract

Spoken dialogue language models (SDLMs) can support spoken interaction by modeling turn-taking behavior and semantic content. However, most existing SDLMs operate on semantic speech tokens, which may limit the modeling of fine-grained acoustic information. Furthermore, it is unclear whether their dialogue understanding benefits diverse speech generation tasks. In this work, we introduce Continuous Spoken Dialogue Language Model (CSDLM), a foundation model that models two-channel spoken dialogue in a continuous latent space. By modeling time-aligned channels through a unified interleaved text-speech stream, CSDLM supports diverse tasks, including both single-channel and two-channel speech generation, without task-specific fine-tuning. We also investigate effective voice-prompt conditioning designs with classifier-free guidance to improve speaker similarity in autoregressive flow-matching speech generation settings. We evaluate CSDLM on a suite of tasks including dialogue continuation, dialogue generation, contextual text-to-speech (TTS), and streaming TTS. Compared with other spoken dialogue models like dGSLM, CSDLM enables fine-grained control over speaking activity while improving the semantic coherence of the generated dialogue by almost 2x. Despite being a streaming model, it matches the lexical and timbre fidelity (in terms of word error rate and speaker similarity, respectively) of non-streaming baselines like VibeVoice and ZipVoice-Dialog, while generating channel-separated dialogue from a reference text. Audio samples are available on the project page: https://anonymous.4open.science/w/temp-3B2C/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.