acceptodds
Under review as a conference paper at ICLR 2027

Transformer-FSQ: Discrete Latents for Extreme KV-Cache Compression

Abstract

Private, personalised language models must run on local consumer devices, where frontier models cannot fit and capability depends on what the model retains in context. However, context memory—the key-value (KV) cache—grows linearly with retained cache entries, quickly dwarfing the model itself: for Qwen3-1.7B-Base, 200,000 retained entries require 21 GiB, 6.6 times more than the model weights. While high token capacity is urgently needed on edge hardware, the KV-cache compression methods we measured fail beyond modest compression ratios. We introduce *Transformer-FSQ*, addressing extreme KV-cache compression by replacing standard key and value projections with a finite scalar quantisation bottleneck. Rather than storing large vectors, the cache retains a packed vector of discrete integer coordinates per retained token per layer, requiring no per-token quantization metadata. Converting a pretrained transformer requires training these projections on 0.25 tokens per original parameter.We evaluate Transformer-FSQ against five baseline families: token eviction, scalar and vector quantisation, low-rank projection, and latent attention on models from 0.6B to 8B parameters. On long-context retrieval (RULER), the methods fall into two regimes. Up to about 5, vector quantisation and low-rank projection stay close to the uncompressed model; from 12 to 97.5, Transformer-FSQ is the most accurate method among our primary evaluated implementations, and leads at extreme compression across scales. Extreme KV-cache compression is therefore a practical way to extend how much context a model on a fixed device can hold.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.