acceptodds
Under review as a conference paper at ICLR 2027

Don't Drop Tokens, Dim Them: Bit Allocation for Multi-Turn KV-Cache Compression

Abstract

Agentic language models spend substantial wall-clock time paused while waiting for tool responses or user input. During these pauses, their KV caches either re- main resident in GPU memory or are managed by a memory store and transferred between storage and the GPU. How to compress the paused cache, and which compression best preserves answer quality over the turns that follow, has no set- tled answer: the two training-free families, token selection and quantization, are both built and evaluated for a single prompt answered immediately. We show that both fail in the paused regime, for opposite reasons. Token selection commits a keep-set before the queries arrive, so methods perfectly recover a planted fact on a single-turn needle collapse on SCBench key–value retrieval. Quantization makes no such commitment but spends its budget uniformly, and on multi-turn retrieval it collapses at 2 bits, which is also the most it can compress. We trace both failures to one cause: precision is uniform and token selection is binary. We therefore propose SOFT-BIT KV. Rather than choosing which tokens survive, it gives every token its own bit-width, fixed once at the pause by rate–distortion allocation under a measured-byte budget. A small MLP scores each cached to- ken from the resuming query’s own per-head attention, and reverse water-filling spends a byte budget over those scores. Average across four models, SOFT-BIT KV retains 92% of full-cache accuracy on multi-turn SCBench at 5.2× compres- sion, compared with 85% for TurboQuant and 72% for the best token-dropping method. At ≈ 11× compression, SOFT-BIT KV retains 99% on single-turn needle retrieval and 65% on long-context RULER, compared with 61% and 56% for the best token-dropping method, and 38% and 23% for TurboQuant. At the system level, it reduces the paused cache of a 98k-token Llama-3.1-8B trajectory by 81% and accelerates host-to-device transfer upon resumption by 5.2×.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.