Deep Low Rank Projector for Low Rank KV Cache Compression
Abstract
Key-value (KV) cache memory grows linearly with context length and becomes a major bottleneck in long-context language model inference. We propose Deep Low-Rank Projector (DLRP), a framework for compressing the KV cache along the head dimension of pretrained attention models. DLRP first trains full-dimensional Deep Linear Projectors (DLPs) on a frozen backbone using a KL objective and a regularizer that encourages low-rank structure. We establish an upper-bound relation between this regularizer and the nuclear norm of the composed projection matrix, providing a tractable surrogate that avoids direct nuclear-norm computation during training. Layer-wise reduced dimensions are then selected from the learned singular-value spectra, and the trained projectors are reduced and fine-tuned for recovery using only a small amount of data. Finally, the resulting DLRPs are folded into the attention projections, leaving no additional projector modules in the decoding path. Across Qwen3 and Llama-3 models, DLRP preserves downstream performance better than the compared hidden-dimension KV-cache compression methods over a range of compression ratios, remains compatible with token-axis compression, and reduces long-context memory usage while increasing feasible batch size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.