acceptodds
Under review as a conference paper at ICLR 2027

Error-Aware Flexible-Length Diffusion Language Model

Abstract

Diffusion language models are trained to generate text by iteratively refining corrupted sequences in a non-autoregressive way. This makes them particularly attractive for error correction tasks. However, standard diffusion language models are usually trained on synthetically corrupted sequences and are limited to a fixed output length, which significantly constrains their performance. We propose a novel Flexible-Length Diffusion Language Model, which addresses the problem of fixed output length by allowing the model not only to substitute tokens during the reverse process but also to insert and delete them. Our model achieves a 21.25% relative improvement over the best USDM configuration on HumanEval single-line code infilling, while maintaining comparable text generation performance. Moreover, we question whether random errors introduced by the standard forward process during training are sufficient for error correction. In particular, we study diffusion models as error correctors in automatic speech recognition (ASR) settings and introduce a framework for using TTS and CTC as a forward process to generate realistic ASR errors and train diffusion models with this task-specific corruption. Additionally, we extend existing methods for using diffusion models for ASR by proposing new decoding methods that not only significantly reduce the gap to autoregressive language models but also achieve a better WER-RTF trade-off. Our Flex ASR-DiffLM model achieves a WER of 3.41% on LibriSpeech dev-other, surpassing the autoregressive language model on several evaluation sets while performing recognition 1.5-2x faster. We publish all training and evaluation recipes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.