Construction Confidence Need Not Be Verification Confidence: Round-Trip Verification and Selective Repair for Masked Diffusion Speech
Abstract
Masked diffusion models can reduce computation by decoding in only a few iterative steps, but in speech generation this low-step regime can leave a minority of outputs with content failures, and which utterances fail can vary across sampling seeds. This suggests a selective strategy of spending extra computation only on outputs likely to fail. The challenge is identifying them. We study this problem in MaskGCT, a zero-shot text-to-speech system whose two stages are masked-diffusion decoders. The first maps text to semantic tokens, a discrete representation of what is said, and the second maps semantic tokens to acoustic tokens that are rendered into speech. The decoder therefore commits tokens, while the failure criterion is defined only after those tokens have been converted into a waveform. We define a content failure as rendered speech whose transcript has a word error rate (WER) above 10% against the target text, that is, speech whose content is unreliable rather than merely imperfect. This distinction gives rise to two kinds of confidence. Construction confidence comes from decoder-derived signals, such as token confidence and entropy, while the sequence is being committed. Verification confidence comes from checking the rendered speech for evidence that the output actually meets that criterion. Used as a failure detector on held-out data at an eight-step decoding budget, construction confidence falls well short of verification. A probing classifier trained on summaries of these signals (the learned probe) reaches AUC , whereas round-trip transcription with a separate small ASR verifier reaches . Round-trip transcription stays above every tested decoder-derived signal at the other decoding budgets we examined, and again in the acoustic stage. It also does so under failure labels from ASR families different from the verifier's, although the margin is smaller. Because a masked-diffusion sequence remains editable, verifier-localized spans can be re-masked and regenerated in place, whereas an autoregressive decoder would have to regenerate everything after them. This selective repair reduces transcript error relative to the unrepaired output across four ASR evaluators. These results show that, when a generator is judged only after its committed tokens have been transformed further, confidence used to construct those tokens need not be sufficient to verify the final output.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.