Variable-Length Non-Autoregressive Speech-to-Speech Translation, With and Without a Pretrained Diffusion Language Model
Abstract
Non-autoregressive generative approaches, such as masked diffusion language models and flow matching, have gained considerable traction in recent years: masked diffusion for text generation and code infilling, and flow matching for image generation and text-to-speech synthesis. Yet they remain rarely explored for speech-to-speech translation. A central obstacle is output length. Source and target utterances differ in length, and the relationship varies across language pairs, whereas in their standard form these models fill a canvas of output positions whose size must be fixed before decoding begins. In this work, we present several techniques for handling variable-length output in non-autoregressive speech-to-speech translation, adapting methods originally developed for other domains, such as code infilling and molecule generation. We study two settings: a large pretrained masked diffusion language model, LLaDA, together with variable-length extensions of it, and a much smaller model based on Branching Flows that does not rely on a large language model. All systems are trained and evaluated on the CVSS-C French–English benchmark. The Branching Flows model reaches 24.19 ASR-BLEU, slightly above the discrete diffusion-based model Dub-S2ST (23.35).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.