TesseraKV: KV Paging via Co-Retrieval
Abstract
Agentic workloads increasingly operate over long contexts, causing KV caches to grow rapidly and making memory bandwidth a major bottleneck during generation. KV paging alleviates this cost by retaining the cache while retrieving only the pages needed at each decoding step. There are two critical challenges in KV paging: how keys are grouped into pages and how pages are scored for retrieval. Existing methods construct pages using heuristics, and score those pages using a separate criterion. These two decisions should instead be derived from a common objective: whether the keys in a page are likely to be retrieved together. We propose TesseraKV, which derives both page construction and page scoring from : keys should share a page when their attention logits tend to vary together under the query distribution of a head. This common principle yields a covariance-weighted geometry for partitioning keys and a spread-aware score for estimating the attention mass of each page. Because the decoding-query distribution is unavailable when the cache is constructed, we further study how to estimate this geometry from queries available beforehand. We find that prefill queries provide a useful but task-dependent signal for decoding-query geometry, while incorporating queries from the downstream instruction substantially improves the co-retrieval structure relevant to partitioning. Across a range of LLMs, long-context tasks, and agentic workloads, TesseraKV consistently improves downstream generation quality over state-of-the-art KV retrieval methods at matched KV budgets. For example, on Claw-Eval agentic tasks, TesseraKV achieves an overall success rate of , improving over the previous state of the art at by points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.