ORG-KV: Finding Optimal Rank and Grouping for Low Rank KV-Cache Compression
Abstract
Low-rank KV-cache compression reduces LLM memory usage by storing compact key and value representations, but introduces a fundamental trade-off between accuracy and efficiency: Higher ranks better preserve accuracy but increase both memory and decoding costs and larger decomposition groups improve accuracy at the cost of decoding computation. Existing approaches typically fix the decomposition granularity in advance and subsequently allocate or search for ranks, potentially excluding more favorable granularity and rank configurations. %from costly training or empirical search. We propose ORG-KV (Optimal Rank and Group KV), a low-rank KV-cache compression framework that jointly finds optimal rank for each key-value projection and decomposition granularity, under explicit KV-cache memory and decoding computation resource constraints. %, without requiring additional training. Specifically, we formulate finding joint rank allocation and grouping granularity as a constrained integer optimization that minimizes estimated compression error under KV-cache memory and decoding FLOPs constraints. Experiments across multiple LLM models demonstrate that ORG-KV consistently outperforms existing training-free low-rank KV-cache compression methods under the same compression budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.