acceptodds
Under review as a conference paper at ICLR 2027

H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache

Abstract

Block diffusion drafters such as DFlash and DSpark reduce drafting latency in speculative decoding by predicting multiple future tokens in parallel. These drafters condition their prediction on the target model by projecting target hidden states at every input position into a separate drafter-side KV cache. This drafter-side cache consumes extra GPU memory per request, reducing maximum concurrency. Maintaining separate caches can also lead to fragmentation or padding in paged serving systems, increasing memory waste. Direct target KV reuse alone eliminates this extra cache but reduces draft quality at later block positions. We propose a hybrid target-context injection method that complements direct target KV reuse with target hidden states only at the last input position, requiring no separate drafter-side KV cache. Based on this design, we propose H-Spec, a hybrid Mamba-attention parallel drafter. Mamba modules are initialized with projected last-token target hidden states, while attention modules reuse target KVs in place. Despite its recurrent formulation, Mamba's parallel scan allows H-Spec to preserve block-parallel drafting. Across three targets and three workloads, H-Spec achieves 5.1-17.3% higher peak throughput across the evaluated concurrency levels and 4.3-24.6% lower average KV utilization than the best baseline for each metric. Despite having no drafter-side KV cache, H-Spec preserves draft quality, improving mean accepted length by 5.0-13.3% over the best baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.