acceptodds
Under review as a conference paper at ICLR 2027

LLC: A Linear Low-Rank Codec for Joint Key-Value Compression Before Token Eviction

Abstract

Token eviction controls KV cache growth but permanently discards states that may become relevant later. Retaining these states in compressed form requires a compact representation that remains useful to attention. We introduce LLC, a closed-form, training-free codec that jointly encodes post-RoPE keys and values across all heads into a shared low-rank code. Fixed linear maps enable direct attention over the codes without reconstructing full-width KV states. We deploy the codec in a two-tier cache: a base eviction policy manages exact residence, while a separately budgeted compressed tier retains demoted tokens with the lowest reconstruction errors. The joint representation improves reconstruction fidelity and yields stronger accuracy–memory trade-offs across models and eviction policies. Across RULER at 8k, 16k, and 32k, LLC approaches dense accuracy with less KV cache memory; at 32k, it stays within 1.5 points of dense across four models using 31–38% of the dense cache. On LongBench, LLC outperforms a prior codec in all eight matched-memory comparisons. Under CUDA Graph replay, an A100 implementation achieves – speedups over matched dense SDPA in single-token, per-layer attention microbenchmarks from 8k to 64k. INT8 quantization of the stored codes further halves their storage with negligible accuracy loss. Together, LLC mitigates irreversible eviction without sacrificing memory efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.