MeanFlow for One-Step Streaming Text-to-Speech
Abstract
A text-to-speech service must start speaking soon after each request arrives while batching many requests on each GPU. Non-autoregressive flow-matching models denoise the whole utterance at every step, so they stream nothing until they finish and their latency grows with the batch; autoregressive models stream, but generating token by token takes many sequential steps or cascades of models. We study a simpler design: a single causal model that generates speech as blocks of continuous codec latents, each in one to three network evaluations, trained from scratch with a MeanFlow objective. We find that the block size trades generation difficulty for speed, that classifier-free guidance must be built into the training target and pushes the output away from real speech, and that a short fine-tuning stage with a critic in a pretrained speech representation improves one-step speaker similarity and quality at every block size at no inference cost. Trained on LibriSpeech, the fine-tuned model achieves near-human intelligibility in one step, produces its first audio in under 20 ms, and delivers more than four times the peak throughput of recent flow-matching models at their fastest intelligible settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.