acceptodds
Under review as a conference paper at ICLR 2027

Improving Parallel Decoding of Block Diffusion Language Models with Autoregressive Generation Order

Abstract

Diffusion language models enable flexible arbitrary-order generation, but existing token sampling methods are mostly designed for early masked diffusion models (MDMs). In this work, we study token sampling for recent block diffusion language models (BDLMs). We show empirically and analytically that these models are naturally more aligned with left-to-right decoding than MDMs. Beyond their AR-pretrained initialization, we attribute this alignment to the block diffusion training objective, which more frequently exposes the model to the left-to-right context used during AR decoding. Based on this observation, we propose Parallel Autoregressive Decoding (PARD), a simple training-free sampling method that preserves left-to-right unmasking structure while allowing parallel token commitment. Experiments show that PARD consistently improves the speed-quality trade-off over existing parallel samplers. Across three recent BDLMs and six benchmarks, it achieves the best average generation quality among parallel samplers. Compared with pure AR decoding, PARD achieves substantial speedups with only a small quality gap.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.