acceptodds
Under review as a conference paper at ICLR 2027

QREAM: KV Cache Compression via Pivoted QR and Explicit Attention Matching

Abstract

Large key-value (KV) caches are a major bottleneck for long-context large language model (LLM) inference, motivating methods that reduce their size while preserving information needed for future queries. Attention matching (AM) has recently emerged as a practical and theoretically grounded framework for controlling attention behavior, with successful applications to prefix tuning and KV-cache compression. However, existing AM-based methods for KV-cache compression separately match the numerator and normalization factor of attention, requiring auxiliary per-key biases that are incompatible with standard KV-cache representations. In this work, we propose QREAM, a bias-free formulation of attention matching that directly matches the normalized attention output rather than its numerator and normalization factor separately. Once the compressed keys are selected, this formulation reduces attention matching to an efficient linear solve over the compressed values, requiring neither training nor modifications to the standard KV-cache representation. To select informative keys, QREAM further introduces a pivoted-QR-based procedure that favors linearly independent attention contributions, providing an efficient alternative to matching-pursuit methods. We evaluate QREAM across multiple Qwen model scales, long-context QA benchmarks, and compression ratios from 5× to 100×. QREAM is competitive with existing attention-matching methods while substantially outperforming token-scoring baselines under aggressive compression, despite its simpler and more deployment-friendly formulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.