acceptodds
Under review as a conference paper at ICLR 2027

LatticeSpec: From Draft Chains to Verification Lattices for Efficient Diffusion Language Model Decoding

Abstract

Masked diffusion language models can predict updates at many positions in parallel, but turning these predictions into faster decoding requires deciding which updates to accept and in what order. Speculative decoding addresses this challenge by proposing updates in advance and verifying them against a target decoder. Yet a draft can contain the right updates in the wrong order: acceptance along a fixed draft chain stops at an ordering mismatch, leaving useful candidates unaccepted. Allowing alternative orders can recover this progress, but their verification costs can erase the benefit. We introduce LatticeSpec, a training-free, plug-and-play extension that selectively relaxes local draft ordering constraints within a small verification budget. It retains the existing drafting procedure and organizes alternative orders into a lattice, sharing verification states whenever different paths commit the same updates. Across 5 tasks and 6 model configurations from the LLaDA, Dream, and SDAR families, LatticeSpec improves end-to-end throughput over the corresponding SSD baselines by 5.7% on average. Separate analyses show that local relaxation recovers updates rejected by chain verification and that state sharing reduces the cost of checking the same alternatives. These results show that recovering more value from existing drafts through limited order flexibility can yield practical decoding gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.