acceptodds
Under review as a conference paper at ICLR 2027

ByteKV: Reusable KV Cache Compression via Bit Allocation and Model Adaptation

Abstract

The key–value (KV) cache accelerates autoregressive decoding in large language models by avoiding redundant context computation, but storing long contexts incurs substantial memory overhead. Reusing a compressed cache across questions requires preserving context information within a tight byte budget before downstream queries are revealed. Here we present ByteKV, a framework combining existing query-blind selection rules with bit allocation and model adaptation. Lowering KV precision enables ByteKV to retain more context positions per byte, while a shared low-rank adapter (LoRA) adapts the model to the sparse, quantized cache through self-distillation from a frozen teacher using a full 16-bit KV cache (all context positions). Across four selection rules evaluated on Llama and Qwen models, 4-bit ByteKV outperforms matched frozen 16-bit baselines on RULER-13, delivering absolute accuracy gains of 17.7 to 58.1 percentage points at a nominal 3% budget relative to full 16-bit KV storage. Integrated with custom fused packed-KV attention and vLLM memory management, ByteKV delivers 1.91× the peak output throughput of full-cache serving with the same adapted weights under concurrent serving. Our code is available at https://anonymous.4open.science/r/ByteKV.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.