acceptodds
Under review as a conference paper at ICLR 2027

When to Trust the Past: Reliability-Calibrated Draft-Tree Orchestration for Speculative Decoding

Abstract

Speculative decoding accelerates autoregressive inference in large language models (LLMs) through a guess-then-verify paradigm, in which multiple candidate continuations are drafted and verified in parallel. Recent methods combine suffix-based retrieval with target-model logits recycled from previous verification steps to reuse historical continuations, thereby expanding candidate coverage. However, such reuse implicitly assumes that historical predictions remain reliable under the current prefix, while poorly transferable candidates occupy limited verification capacity and crowd out more promising alternatives when this assumption fails, an issue which is further exacerbated as history accumulates, since the same token recurs under increasingly diverse prefixes. Towards this end, we propose a novel framework termed Reliability-Calibrated Draft-Tree Orchestration for speculative decoding (RECAP). The core of RECAP lies in estimating prefix-dependent reliability of historical predictions and then using this calibrated evidence to jointly determine candidate composition and the allocation of the verification budget. In particular, we first construct a multi-context verification memory that stores target-model predictions together with their corresponding prefix contexts. Subsequently, we introduce a lightweight prefix-conditioned scorer that estimates the transferability of each memory entry to the current decoding state and aggregates reliable historical evidence. Based on the resulting evidence, RECAP combines retrieved continuations with logit-based candidates under a shared verification budget and expands the draft tree best-first along the paths of highest cumulative reliability. Together, these components transform draft-tree construction from a fixed heuristic into an orchestration process that adapts both candidate sources and tree shape to the available evidence. Extensive experiments show that RECAP consistently improves decoding efficiency over strong speculative-decoding baselines while preserving generation quality. The implementation of RECAP is available at https://anonymous.4open.science/r/RECAP7C2E.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.