DiffuSpeech: Bidirectional Co-Refinement of Reasoning and Speech via Unified Speech-Text Diffusion
Abstract
Recent autoregressive (AR) speech language models have begun incorporating intermediate text reasoning before producing spoken responses, mirroring chain-of-thought in text LLMs. However, the AR paradigm enforces a strictly sequential pipeline: reasoning tokens are committed left-to-right and cannot be revised once subsequent speech tokens are emitted, leaving no mechanism for the spoken response to in turn shape the reasoning that produced it. We argue that Bidirectional Co-Refinement of reasoning and speech—where thought guides speech and emerging speech refines thought—is a more powerful paradigm, and one that is uniquely enabled by unified speech-text diffusion. We introduce , the first diffusion-based speech-text language model supporting both interleaved speech-text understanding and generation, which jointly denoises text reasoning traces and tokenized speech under a single masked diffusion framework with modality-specific masking schedules. At every denoising step, the model conditions on all currently unmasked tokens across both modalities, allowing reasoning and speech to mutually refine each other throughout generation. To enable this paradigm, we further construct , a speech QA dataset with paired text reasoning traces (26K samples, 319 hours). On the three speech QA benchmarks supporting S S evaluation, attains state-of-the-art speech-to-speech accuracy among 10B speech-generative models while preserving the base LLM's language-understanding ability and delivering competitive TTS quality. Most diagnostically, on the same matched benchmarks, it combines the highest S S accuracy with a smaller S T-to-S S degradation than high-performing AR baselines such as Qwen2.5-Omni, MinMo, and GLM-4-Voice (9.1 vs. 10.7–14.3 points)—evidence for Bidirectional Co-Refinement—and a strictly controlled 1.7B ablation plus an 8B replication show that the advantage is architectural rather than an artifact of scale or reasoning data alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.