TapKV: Accelerating Long-Horizon Speculative Decoding with Target-Aware Draft KV Caches Optimization
Abstract
The speedup of speculative decoding hinges on generating draft proposals at low cost while maintaining high draft–target agreement. In long-horizon speculative decoding, where generation spans long decoding sequences, however, the growing draft KV cache incurs substantial memory traffic and increases drafting latency. KV-cache pruning provides a natural opportunity to reduce this overhead and accelerate drafting. However, existing attention-based importance metrics primarily identify influential KV entries for draft model performance, without explicitly capturing how their removal affects draft–target agreement and, ultimately, the speculative acceptance rate. We introduce TapKV, a draft-cache pruning method. We derive an Attention-output Taylor (AOT) score that combines the deletion-induced attention-output change with the loss gradient, capturing both value removal and attention renormalization. AOT guides cache-entry selection and global budget allocation across layers, while a paged-attention kernel efficiently supports the resulting unequal cache lengths. Experiments across diverse workloads show that TapKV achieves higher acceptance lengths and speedups than existing baselines, delivering an average speedup of –.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.