acceptodds
Under review as a conference paper at ICLR 2027

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

Abstract

Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read at every decode step. We find that attention keys are approximately low-rank within pages. A single low-rank projection shared across pages can miss page-specific directions; fitting a basis to each page better identifies the pages receiving the most attention at comparable stored selector cost. LOCKS stores a rank- spectral summary per page, reconstructs its within-page logits, and selects pages by log-sum-exp mass without reading candidate keys or values. It stays within about a point of FullKV on LongBench-v1, tracks the read-every-key exact-LSE oracle on RULER down to the smallest budgets, and retains quality furthest under tight budgets on AIME26 and MATH-500. At a -token budget it matches FullKV aggregate quality beyond K context while attending about % of tokens. Across ranks -, summaries use - of full-KV bytes. On GH200 with GPU-resident KV, LOCKS reduces complete decode-step time by at K context. With full KV offloaded to Grace memory, it reaches - the faster dense backend's aggregate throughput at K-K by serving larger batches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.