DChord: Schema-Aware Speculative Decoding for Structured Generation
Abstract
Agents and embodied systems increasingly rely on large language models to produce structured generation for information extraction, tool invocation, and action commands. Recurring workloads are widely applied in practical scenarios, where requests from the same task share predefined field names, field order, and serialization, while the model predicts short, input-dependent field values. Recent work has explored speculative decoding to accelerate these structured generation workloads. However, these methods typically draft the complete serialized token sequence, allocating prediction positions to schema-determined structure and risking avoidable rejections on tokens that need not be predicted. We present DChord, a schema-aware speculative decoding framework with two components: PairDraft and PairCut. PairDraft concentrates prediction exclusively on field values and end markers, while a pair state machine supplies schema-determined structure to expand a compact value-only prediction into a complete draft. For concurrent serving, PairCut treats schema-determined structure as certain in its verification decisions, increasing acceptance length with negligible additional target latency. We evaluate DChord on Qwen3.5-4B, Qwen3.6-35B-A3B, and Gemma4-12B across four schema-determined tasks spanning multimodal and text-only inputs, comparing against autoregressive decoding (AR) and four recent speculative decoding methods. At concurrency 16, DChord achieves up to throughput speedup over AR and up to the throughput of the strongest competing speculative method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.