Does a Grammar-Constrained Speculative Verifier Sample the Conditional Law? Fidelity Bounds, Numerical Paths and Exact Acquisition
Abstract
A speculative verifier samples whatever target it is supplied, so grammar-constrained speculative decoding samples the model's grammar-conditional law only if the verifier receives the future-validity-corrected target of grammar-aligned decoding. We measure what an implemented stochastic verifier actually samples once that target is supplied, on finite canonical JSON events whose conditional law is enumerated exactly. In a Qwen3-8B/DFlash block verifier, simultaneous 95% bounds place the terminal law within 0.028 total variation (TV) of the conditional law on four events, removing at least 88% of each masking gap; a registered chain-draft verifier and a second target, Qwen3-14B, reach the same conclusion, and a crossed dtype × execution-path family places every rejection in the bf16 native block forward. We show that the masking gap is the relative dispersion of the legal mass along sampled paths and cannot exceed one minus the event's probability. On 25 matched events, stating the schema in the prompt lowers the median gap from 0.28–0.48 to 0.009–0.051, and across instances the gap follows the visit-weighted dispersion of future validity. The correction itself costs the verifier 0.03 per-position acceptance. To acquire the target exactly, we prove that bounds from acquired rows alone certify a resolved-leaf sampler to within a factor two of the best terminal guarantee those rows allow once valid mass is resolved, and that draft span proposals sample the conditional law for every reachable cache state; adaptive exact samplers need as little as 0.019 of full enumeration's wall clock, and span proposals save target forwards when each forward recomputes a long prompt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.