acceptodds
Under review as a conference paper at ICLR 2027

Grammar Reachability for Lossless LM-Head Compression and Deterministic Token Injection in Structured Generation

Abstract

Schema-constrained generation offers two complementary deployment optimisations. GRVP (Grammar-Reachable Vocabulary Pruning) identifies the unique minimum LM-head row set preserving the grammar-masked output distribution (Theorem 4.1). It removes 98% of rows on automotive Thrift IDL and 99.9% on public VehicleWorld, with ΔEM=0.0000 across nine configurations and three model families; tied and mixed-type models use masks without storage savings. GGSD (Grammar-Guided Structured Decoding) injects grammar-forced tokens without a forward pass and reproduces per-state constrained greedy decoding for every block budget (Theorem 5.4). Reported speedups are 1.8× on VehicleWorld, 5.86× on xLAM and 64× on automotive (GPU-synchronised median; up to 86× on a favourable subset), with deterministic fractions of 46%, 84% and 96.5% respectively. Both methods are post-training and require no training data. Frequency pruning needs Ω(|L(GS)|) samples in the worst case: 26% of reachable automotive tokens are absent from training, whereas xLAM deterministic-position coverage reaches 96%. Grammar analysis thus separates structural reachability from corpus frequency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.