ZENO: Zone Expanding Decoding Operator for Diffusion Language Models
Abstract
Diffusion Large Language Models (dLLMs) generate text by unmasking tokens from anywhere on the generation canvas. This maximizes generation flexibility but minimizes context buildup by forcing models to commit to premature conclusions before establishing reasoning chains. This lack of directional grounding leads to poorly formed trajectories, high inference latency and rigid denoising schedules. We introduce ZENO (Zone Expanding decodiNg Operator), a training-free sampler that transitions from autoregressive grounding to parallel diffusion within a single generation instance at inference time. ZENO enforces early directional priors by restricting generation to a geometrically expanding prefix window, then shifts to a fixed block once sufficient context crystallizes. Evaluated on six standard benchmarks, across base and instruction-tuned checkpoints of LLaDA and Dream families, ZENO achieves up to a and speedup in per-query latency, with token throughput gains of up to and and cuts reverse diffusion steps by up to %. ZENO also improves generation quality, yielding a mean accuracy gain of up to points over vanilla decoding. Furthermore, enabling KV caching reduces per-query latency by up to with small accuracy trade-offs, establishing ZENO as a fast and efficient decoding strategy for dLLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.