Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding
Abstract
Large language models (LLMs) are increasingly applied to very long inputs, such as whole documents, codebases, and conversations that run for many turns. The key-value (KV) cache stores the attention keys and values of all past tokens. Because the cache grows with the context length and is re-read in full at every generated token, it dominates the inference memory at long context. Most existing methods reduce this cost by quantizing the cache: storing every value at the same low precision. These methods reach about two bits per value, but they rarely go below two bits, because accuracy drops sharply at that point. The reason is that two bits allow only four levels per value, which are too few for a cache with outlier channels, which is known as the 2-bit cliff. In our study, we find that once we project the cache into a latent space that removes the correlations between channels, the contributions of the different channels of the KV cache to the attention output are highly uneven. And we find that spending the bit budget on those few important channels is far more accurate than spreading the bit budget evenly. Guided by these observations, we develop SPECTRA, a training-free compression method for the KV cache that requires no change to the model. SPECTRA transforms the KV cache into a latent space in which the channels are ordered by their contribution to the attention output, and then allocates the bit budget across these ordered channels. On long-context benchmarks and three instruction-tuned models, every uniform quantization methods that we compare against has fallen below the 2-bit cliff at 8 compression rate, and none of these methods has a configuration beyond 9. SPECTRA matches the uncompressed model at 8 on Llama-3.1-8B, stays within 1 point of that accuracy up to 11, and recovers every planted needle in needle-in-a-haystack at 12.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.