SealKV: Accurate Low-Bit KV Compression for Faster LLM Inference
Abstract
Long-context reasoning and agentic workloads have made the key–value (KV) cache the dominant cost of LLM serving: it is re-read at every decoding step and grows with every generated token, bounding both decoding speed and concurrency. Compressing it pays off only if it combines high quality, low memory use, and fast serving, yet existing quantization methods sacrifice at least one: they require offline calibration, lose accuracy at low bit widths, or save memory without a measured speedup. We introduce SealKV, the first KV-cache compression method that achieves all three, by bringing a trellis-coding quantizer to a cache that grows continuously. Our key insight is that quantization can take place in sealed blocks, each coded only once all its tokens have arrived; sealing decouples trellis encoding from the newest token and exposes a subtractable per-channel offset, shared by neighboring tokens, that rotation alone cannot remove; subtracting it also lifts three existing quantizers by up to 32 points on GPT-OSS-120B retrieval. We also observe that a shorter trellis is also affordable online at little loss: the code's quality depends far less on its window size than its encoding cost does. Co-designed fused dequantization–attention kernels read the codes directly and reconstruct values in registers without writing them back to memory. Without calibration, SealKV is effectively lossless at three bits and above and state-of-the-art at two, where it shrinks KV cache memory by while losing as little as 0.37 points on Qwen3-14B and outscoring all evaluated baselines. Integrated end to end into vLLM, SealKV raises peak decode throughput at 64K context to that of BF16 at 4 bits and at 2 bits, ahead of all evaluated baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.