Caduceus: Parallel Speculative Decoding with Two Simple Drafters Intertwined
Abstract
Speculative decoding can reduce target-model passes by verifying multiple draft tokens at once, but its end-to-end speedup depends on balancing proposal quality against drafting and verification cost. Small language models offer high-quality drafts at the cost of latency, whereas n-gram retrieval is fast but depends on close matches. We introduce Caduceus, a training-free speculative decoding framework that combines these complementary drafters around a persistent draft tree. Within the compound drafter, n-gram retrieval and small-model expansion alternate: retrieval adds candidate paths to the persistent tree, and the small model scores and extends the tree to produce new anchors for subsequent retrieval. Target observations also update a prompt-local online n-gram store. At runtime, the drafter continues growing the tree while the target model verifies a snapshot, and compatible branches persist across verification rounds. Across four target models from 8B to 70B, Caduceus achieves 2.13–3.70× speedup on Spec-Bench and 2.27–3.86× on HumanEval over target-only autoregressive decoding, attaining the highest full-benchmark throughput among the baselines reported for each of the eight benchmark–model settings. On repeated-sampling SpecTTS-Bench with a Llama-70B target, Caduceus achieves 3.99× speedup on the first rollout and 4.39× on the third.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.