acceptodds
Under review as a conference paper at ICLR 2027

The Prefix Is Already Cached: Reusing Target KV for Speculative Decoding

Abstract

A compact drafter can make several proposals before verifying against a single target model. Head-based speculative drafters use the target hidden states to build their proposals, to reconstruct a verified-prefix representation. Target model inference already cached the prefix representation in its own native key-value (KV) states. In same-query experiments, we show that a system with duplicate hidden KV states can change how future tokens are interacting with earlier context, and more importantly, how these states can cause inaccuracies in verified prefix readouts when queried by the target itself. Particularly we find that when a continuation visits information receiving low attention at the current token location, readouts can be subject to larger inaccuracies in these hidden KV state reconstructions, causing non-negligible errors in rejecting later steps. Based on these observations, we introduce OpSpec, which removes the need for prefix reconstruction, where each branch simply queries the target's native KV states without any sequence linear drafter state over the verified token sequence. Operator Readout Distillation is used to distill the target model's prefix readouts with queries over these KV states, and a fused attention kernel allows for caching pages without copying. OpSpec extends accepted length at 7 benchmarks on Qwen3-8B at the reported tree budget. Much larger gains are observed on tasks that require a larger re-visiting of earlier context. At 32K context length, OpSpec saves 96.5% of incremental drafter state, allowing 20% more maximum batch size, and 9.7% more throughput at each method's maximum feasible batch.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.