acceptodds
Under review as a conference paper at ICLR 2027

CaDiS: Causal Diffusion Transformer for Streaming Speech Generation

Abstract

Diffusion and flow-matching models have advanced high-fidelity speech generation. However, their bidirectional attention requires generating a full utterance before playback, and iterative sampling adds further delay. Most streaming systems therefore rely on autoregressive generators and use diffusion only for acoustic decoding. In this paper, we ask instead whether a pretrained offline diffusion model can generate speech chunk by chunk in real time. We present **CaDiS**, an offline-to-streaming framework built on a causal diffusion transformer that decouples the architectural and acceleration challenges into a three-stage recipe: autoregressive diffusion adaptation finetunes the bidirectional model with block-causal attention to generate each chunk from past speech alone; causal consistency distillation reduces sampling to four steps; and distribution matching refinement restores teacher-level quality. We adapt CaDiS to two streaming tasks, text-to-speech (TTS) and voice conversion (VC), with dedicated task-specific efforts: a normalized positional encoding for asynchronously aligned text in TTS, and content-vector distillation plus reference-anchored positional re-encoding to address the voice consistency and artifact challenges in VC. We evaluate CaDiS on objective and subjective metrics across natural human and expressive anime voices. CaDiS preserves offline-teacher quality while streaming: its TTS model attains a ms first-packet latency and real-time factor on an A100 GPU, while its VC model achieves the best intelligibility and perceptual quality among streaming baselines with no look-ahead. Moreover, CaDiS attains among the best overall streaming efficiency profiles, demonstrating the potential of diffusion speech generators in streaming scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.