Low-Rank Key Value Attention
Abstract
The key-value (KV) cache is a primary memory bottleneck in Transformers. We propose Low-Rank Key-Value (LRKV) attention, which reduces KV-cache memory by exploiting redundancy across attention heads. Each layer pairs a shared KV projection with low-rank, head-specific residuals, interpolating continuously between full sharing and per-head independence. Across pretrained models from 128M to 6.3B parameters, LRKV has the lowest observed held-out loss among standard MHA, MQA/GQA, and MLA at each method's operating point while using 48–67% of MHA's cache, and it remains below full-cache MHA with as little as 8.3% of the cache at 128M, 2.5B, and 6.3B. At 128M and 100B tokens, three-seed means give LRKV lower held-out loss than MLA at every matched cache budget (8.3%, 25%, and 50%). At 2.5B, LRKV reaches each baseline's final quality in 18–30% fewer optimizer steps. Trained at 8K context, 512M LRKV models attain the lowest long-context loss and lead every baseline on RULER at 32K, four times the training length. Its exact decode kernel evaluates the factored representation end to end without storing reconstructed per-head K/V. On a B200 at 2.5B, LRKV serves 8.5–8.9 MHA's resident tokens at every context and decodes at 103–111% of MHA's batch-1 throughput from 1.2B to 30B. After fixed-recipe midtraining, LRKV has the highest observed combined downstream accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.