LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
Abstract
Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-specific directions; fitting a basis to each page better identifies the pages receiving the most attention at comparable stored selector cost. LOCKS stores a rank- spectral summary per page, reconstructs its within-page logits, and selects pages by log-sum-exp mass without reading candidate keys or values. It stays within about a point of FullKV on LongBench-v1, tracks the read-every-key exact-LSE oracle on RULER down to the smallest budgets, and retains quality furthest under tight budgets on AIME26 and MATH-500. At a -token budget it matches FullKV aggregate quality beyond K context while attending about % of tokens. Across ranks -, summaries use - of full-KV bytes. On GH200 with GPU-resident KV, LOCKS reduces complete decode-step time by at K context. With full KV offloaded to Grace memory, it reaches - the faster dense backend's aggregate throughput at K-K by serving larger batches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.