acceptodds
Under review as a conference paper at ICLR 2027

Latent-Plan Block Diffusion: Sampling a Block-Level Intent for Few-Step Parallel Decoding

Abstract

Block diffusion language models such as BD3-LM use a masked-diffusion denoiser to decode a block of tokens in parallel while keeping the key–value (KV) cache of autoregressive models, but every denoising step samples the block from a product of per-token marginals. We show that the least error of drawing a block in one such step is its conditional total correlation given the context, a cost that unmasking schedules can only trade for more steps. *Latent-Plan Block Diffusion* (LP-BD) supplies the missing shared variable: before a block is decoded, a small conditional flow samples a low-dimensional *plan* given the generated prefix, and the denoiser decodes the block conditioned on it. The plan is learned from the denoiser's own loss, without supervision, and LP-BD equals BD3-LM at initialization, so it can be added to a pretrained model. Trained from scratch on TinyStories with 16-token blocks, LP-BD lowers the negative log-likelihood of one-step samples under an autoregressive judge by 31% and reaches 16-step BD3-LM quality with 8 steps. The gains are similar at 28M and 85M parameters. Added to the released 170M BD3-LM checkpoints on OpenWebText, LP-BD lowers generative perplexity at matched entropy by 14–55% at every step budget of a fixed-step sampler, and by 38–66% with guidance. The plan acts as a sampled *intent* shared by the block rather than a prediction from the context, and its gains require steps that decode several tokens jointly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.