Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
Abstract
Large language model (LLM) generation remains expensive because autoregressive decoding requires a model call for each new token. Speculative decoding reduces this cost by drafting multiple tokens and verifying them with the target model in one step, but its speedup depends on how many draft tokens are accepted. Parameter-free draft sources, such as suffix caches, can propose long continuations at low cost in structured and agentic workloads, yet their payoff varies across decoding steps. When a promising cache match yields only a short accepted prefix, most of the verification work is wasted. We propose Hybrid Verified Decoding, which uses a payoff predictor to estimate the accepted length of a cache draft before verification and then decides whether to verify the cache draft or fall back to a model-based drafter. Through experiments spanning three LLMs and sixteen datasets, we show that Hybrid Verified Decoding is especially effective on agentic workflows, where it outperforms EAGLE3 in every setting with speedups of 1.4× to 5.5×. Our analysis demonstrates that high-payoff cache drafts are rare and that the payoff predictor identifies most of them. We further show that the predictor transfers to unseen datasets and other target models without retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.