acceptodds
Under review as a conference paper at ICLR 2027

LARK: Low-Rank Archived Retention of KV Cache via Linear Expected-Attention Scoring

Abstract

Long chain-of-thought (CoT) reasoning generates tens of thousands of tokens, making KV cache the primary memory and bandwidth bottleneck. Token eviction irreversibly discards history; its accuracy hinges on ranking keys by attention from queries that do not yet exist, which content predicts only in their frame, reached by an inverse rotation per candidate. We propose LARK, a KV cache compression system built to archive rather than discard. LARK comprises: (1)LEAP, a linear scoring operator strictly equivalent to pre-RoPE expected-query phase scoring yet executing as a single projection on cached post-RoPE keys, with cost independent of offset and query-sample counts; (2)LACE, a low-rank archival pipeline whose evicted tokens are compressed into factors that leave the model cache yet keep participating in attention, re-scored at segment level by the same form; (3)a systems co-design enabling CUDA-graph-compatible variable-length continuous batching. Across four models and four benchmarks at a k generation limit, LARK retains accuracy far above eviction-based baselines as the budget tightens. At it reaches near-lossless accuracy with 5.0% of the dense device KV memory, with 2.0-3.3x iso-batch throughput.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.