SuffixSpec: Hierarchical Context-Aware Memory for Training-Free Speculative Decoding
Abstract
Speculative decoding accelerates large language model (LLM) inference by drafting candidate tokens and verifying them in one forward pass. Training-free drafters avoid training a draft model for every target model, but they either copy text from the prompt, the generated history, or an external corpus, or index the target model's own predictions by a single token, and none of them uses the predictions that the target model already computes at every prompt position during prefill. We cast drafting as retrieval from a memory of these predictions, organized along two axes: the granularity at which contexts are keyed and the lifetime of stored evidence. Along these axes we build SuffixSpec, a hierarchical memory whose suffix-local, token-local, and token-global layers are queried in order of decreasing specificity; a prefill bootstrap populates it before the first decoding step, and a dual-timescale commit updates the persistent layer once per request with decode-time predictions only. On seven tasks and three models from 8B to 32B parameters, SuffixSpec achieves up to speedup over autoregressive decoding and improves over the strongest training-free baseline by up to . In a controlled comparison inside the SGLang serving engine on Qwen3-8B with 8K-token contexts, it outperforms the trained drafters EAGLE-3 and DFlash under greedy ( vs. and ) and stochastic decoding, while keeping about 10 MB of draft state in host memory instead of 2 to 5 GiB of extra GPU memory for the drafters. The code is available at https://anonymous.4open.science/r/suffixspec-8075.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.