StrideKV: Accelerating Batched Long-Context Decoding via Periodic Reuse of Retrieved KV Cache
Abstract
Query-aware KV cache retrieval enables long-context inference by keeping the full KV cache in CPU memory and maintaining only a small GPU-resident working set of entries relevant to each query, but repeatedly selecting and loading these entries can dominate decoding time. We study how often this retrieval needs to occur in synchronous batched decoding. A natural approach is adaptive: reuse the previous working set and perform a correction, an immediate retrieval for the current query, only when the query changes substantially. However, when requests in a batch execute in lockstep, different requests need corrections at different layers and steps, so the batch as a whole must perform corrections far more often than any single request; as the batch grows, the throughput gain of such adaptive schemes vanishes and eventually turns negative. We also observe that queries at adjacent steps select largely overlapping tokens, and that retrieving only every other step yields the same top-1 token as per-step retrieval at 99.4% of decoding steps. Motivated by these findings, we introduce StrideKV, which retrieves relevant KV entries only once every few decoding steps, with all requests in the batch following the same schedule. When retrieval happens, the entries are still selected based on the current query; in the steps between retrievals, the model simply reuses the entries already on the GPU, avoiding both selection and CPU-to-GPU transfer. Across multiple models and context lengths, StrideKV improves decoding throughput on an A100 GPU by up to 66% over ShadowKV, a strong KV offloading method that retrieves at every step, while keeping long-context accuracy close to that of ShadowKV. These results show that StrideKV is a simple and effective way to accelerate batched long-context decoding with KV cache offloading.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.