Select Twice, Freeze Once: Jacobian-Based KV Cache Selection for Multi-Turn and Long-Context LLM Serving
Abstract
The long-term deployment of multi-turn LLM systems depends on reliable and efficient handling of ever-extending context windows. As the KV cache scales linearly with sequence length, it has become a primary bottleneck for memory and throughput. This overhead is particularly detrimental to multi-turn systems, which generate a high volume of prefill and decode calls and are therefore disproportionately sensitive to the overhead of low-rank approximation techniques. Existing KV cache selection methods rely on the attention map as the primary selection signal. We find that the attention-output Jacobian motivates considering local sensitivity beyond individual attention weights when selecting entries for reuse. We introduce PrescienceKV, which combines a linear surrogate derived from a quadratic scoring proxy with windowed attention-based reranking. The surrogate constructs candidates, while reranking aggregates local value sensitivity over recent prompt queries and the first decode query to select the final top-k historical subset. Both stages run at the first decode step of each turn over the prefilled history, after which the subset is frozen and all newly generated KV entries are appended without further selection. We evaluate PrescienceKV on RULER and LongBench across five models for accuracy and efficiency. PrescienceKV reduces the selected historical KV budget by over 96% at 128K context while maintaining accuracy comparable to full attention. At capacity, PrescienceKV sustains a decode batch of 16-20 across context lengths up to 124K, where full attention's maximum batch drops to 2, yielding decode-throughput speedups up to 6.01x. In the separate per-token latency comparison at 128K, it reduces latency by 33% relative to full attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.