DConstrain: Compiling Context-Parameterized Transitions for Diffusion Language Model Constrained Decoding
Abstract
Diffusion language models (DLMs) enable parallel token generation, but structured-output constraints are difficult to enforce because token choices at unresolved positions can depend on one another. Automata-based DLM constrained decoding resolves this interdependence through joint inference, which requires exact token-labeled successors for many hypothetical parser configurations rather than validity information for a single realized prefix. Existing finite-state pipelines bind runtime parsing contexts into explicit states and materialize vocabulary-level transitions, duplicating reusable grammar logic across contexts. We present DConstrain, which instead compiles context-parameterized transition programs: token consumption is specialized to local grammar states and tokenizer tokens, while the runtime stack and parsing environment remain unbound. Warm compilation executes these reusable programs only for reached configurations, recovering the exact successors required by DLM joint inference. Token-effect grouping further shares successor construction across equivalent tokens. On BFCL with LLaDA2.2, DConstrain compiles 100% of schemas under arbitrary property order versus 71.3% for an optimized finite-state baseline, with a paired cold-compilation speedup. It also reduces constrained-decoding overhead by while preserving 100% schema validity and matched task accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.