Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Abstract
Efficient zero-shot voice cloning remains challenging: autoregressive systems are slow, while many non-autoregressive systems still require reference transcripts during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Tacit-TTS replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10 times faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning also supports cross-lingual and non-lexical references. We validate these settings using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail because of unreliable ASR transcripts, making reliable voice cloning difficult.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.