OrthoKV: Extreme KV Cache Compression with Exact Orthogonal Sparse Coding
Abstract
The KV cache is a major memory bottleneck in long-context inference. Sparse coding is attractive for extreme KV cache compression, but existing approaches rely on iterative sparse approximation over large overcomplete dictionaries, introducing substantial encoding overhead. We introduce OrthoKV, which replaces layer-wise overcomplete dictionaries with complete orthogonal dictionaries specialized to individual KV heads. Per-head specialization helps compensate for the reduced representational capacity of complete dictionaries, while orthogonality makes sparse coding exact and non-iterative. The best -term approximation under squared reconstruction error is obtained by a single transform followed by selecting the coefficients with the largest absolute values. The same result extends across heads. With a total budget of coefficients, one simply keeps the largest coefficients in absolute value across all heads. We learn the dictionaries with a sparsity-promoting objective that does not depend on the target sparsity, so the same dictionaries can be used at different compression ratios without retraining. Complete dictionaries also enable smaller coefficient indices and more compact sparse codes. In the extreme compression regime, OrthoKV uses less KV cache memory than Lexico while matching or improving its quality in most evaluated settings and substantially reduces sparse coding latency. Compared with full KV caching, OrthoKV achieves higher throughput at larger batch sizes and supports batches that would otherwise run out of memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.