acceptodds
Under review as a conference paper at ICLR 2027

Deep Low Rank Projector for Low Rank KV Cache Compression

Abstract

Key-value (KV) cache memory grows linearly with context length and becomes a major bottleneck in long-context language model inference. We propose Deep Low-Rank Projector (DLRP), a framework for compressing the KV cache along the head dimension of pretrained attention models. DLRP first trains full-dimensional Deep Linear Projectors (DLPs) on a frozen backbone using a KL objective and a regularizer that encourages low-rank structure. We establish an upper-bound relation between this regularizer and the nuclear norm of the composed projection matrix, providing a tractable surrogate that avoids direct nuclear-norm computation during training. Layer-wise reduced dimensions are then selected from the learned singular-value spectra, and the trained projectors are reduced and fine-tuned for recovery using only a small amount of data. Finally, the resulting DLRPs are folded into the attention projections, leaving no additional projector modules in the decoding path. Across Qwen3 and Llama-3 models, DLRP preserves downstream performance better than the compared hidden-dimension KV-cache compression methods over a range of compression ratios, remains compatible with token-axis compression, and reduces long-context memory usage while increasing feasible batch size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.