CoFeatKV: Coalesced-Feature Low-Rank Compression for Post-RoPE KV Caches
Abstract
During long-context inference, the KV cache grows linearly with the sequence length and becomes a major source of memory usage and data movement. Existing low-rank key compression faces a trade-off: compressing pre-RoPE keys typically preserves a stronger low-rank structure, but decoding requires key reconstruction followed by RoPE; compressing post-RoPE keys avoids repeated positional rotation, but their weaker low-rank structure often leads to larger compression error. We propose CoFeatKV, a training-free method for post-RoPE key compression. Within each token page, CoFeatKV first groups key features with similar structure, builds a shared representation for each group, and then applies low-rank compression to the grouped representation. The compressed cache stores only the shared representatives, feature coefficients, and group assignments. During decoding, the query is mapped directly into the compressed group space to compute attention scores, avoiding both full reconstruction of historical keys and repeated RoPE on cached keys. CoFeatKV uses page-local updates to support autoregressive decoding while retaining all historical tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.