acceptodds
Under review as a conference paper at ICLR 2027

From Grammar Deltas to Token Masks: Incremental Constrained Decoding under Runtime Grammar Evolution

Abstract

Grammar-constrained decoding enables large language models to generate outputs that satisfy formal constraints, but in code and domain-specific language generation, the relevant constraints are not always fully available before decoding and may emerge progressively from new declarations, inferred type information, or runtime analysis feedback. Existing structured-decoding engines efficiently handle fixed, predeclared, or previously seen constraint structures; however, when a first-seen grammar update changes the parser state induced by an already committed prefix, the straightforward solution reconstructs the constraint state under the updated grammar, repeating parsing work already performed over the prefix. We introduce GramDelta, an incremental framework for constrained decoding with append-only runtime grammar updates. GramDelta maintains parse state over the committed prefix and propagates only changes induced by newly added grammar rules through historical dependencies, avoiding prefix replay. It further incrementally maintains next token validity from these parser state changes, reusing an unchanged mask or rechecking only affected tokenizer-trie regions when the impact is localized. GramDelta therefore preserves the result of full reconstruction while making update cost sensitive to the actual impact of a grammar change. Experiments across four model configurations show that GramDelta reduces median first-seen grammar-update latency by – relative to XGrammar and reduces the additional TPOT overhead over unconstrained serving by –. These results enable efficient constrained decoding when constraints progressively emerge and evolve with the generated prefix.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.