SemanticTree: Grammar-Aware Tree Allocation for Speculative Decoding
Abstract
Tree-based speculative decoding allocates candidate tokens before target verification. In structured generation, grammar checks come too late to reallocate this budget: 73–77% of non-root slots in EAGLE3 trees contain grammar-invalid candidates. We call this semantic waste. SemanticTree excludes unreachable candidates during tree allocation while preserving the proposer's scores and budget. A parser-state mask cache shares work across branches and rounds, and verification reuses the masks obtained during drafting. We integrate SemanticTree with an autoregressive drafter, EAGLE3, and a block drafter, DDTree. Across 35 drafter–target–workload settings at batch size one, verification rounds fall by 16.4–29.8% with EAGLE3 and 17.2–25.8% with DDTree; decode throughput rises by 19.5–40.3% and 15.5–25.6%, respectively. A cache ablation preserves all round savings but reduces the throughput gain from 33% to 7% on Qwen3-4B/Nexus with EAGLE3, showing that recovering semantic waste requires both effective allocation and low grammar-processing cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.