acceptodds
Under review as a conference paper at ICLR 2027

ORG-KV: Finding Optimal Rank and Grouping for Low Rank KV-Cache Compression

Abstract

Low-rank KV-cache compression reduces LLM memory usage by storing compact key and value representations, but introduces a fundamental trade-off between accuracy and efficiency: Higher ranks better preserve accuracy but increase both memory and decoding costs and larger decomposition groups improve accuracy at the cost of decoding computation. Existing approaches typically fix the decomposition granularity in advance and subsequently allocate or search for ranks, potentially excluding more favorable granularity and rank configurations. %from costly training or empirical search. We propose ORG-KV (Optimal Rank and Group KV), a low-rank KV-cache compression framework that jointly finds optimal rank for each key-value projection and decomposition granularity, under explicit KV-cache memory and decoding computation resource constraints. %, without requiring additional training. Specifically, we formulate finding joint rank allocation and grouping granularity as a constrained integer optimization that minimizes estimated compression error under KV-cache memory and decoding FLOPs constraints. Experiments across multiple LLM models demonstrate that ORG-KV consistently outperforms existing training-free low-rank KV-cache compression methods under the same compression budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.