From Memory Savings to Throughput Gains: Serving-Native KV Compression for Sparse-Attention LLMs
Abstract
Long-context LLM serving is increasingly bottlenecked by the growing memory and bandwidth demands of the KV cache. Sparse attention alleviates the KV access cost, but still leaves the full KV cache resident, limiting serving concurrency. Existing KV compression methods primarily target compact storage or optimized kernels, rather than the execution and lifecycle of compressed KV in online serving, where compression timing, efficient attention over compressed states, and cache management directly affect performance. We present Duet, a serving-native KV compression framework that turns KV memory savings into practical serving gains while preserving the underlying sparse-attention selection semantics. Duet first develops an inference-friendly low-rank representation that exploits cross-head redundancy and removes per-token reconstruction from sparse attention. Building on this, a mixed native-latent attention operator directly consumes selected KV blocks in either form without materializing compressed history. Finally, Duet integrates compression into the serving runtime through dual-representation KV management, watermark-triggered batched compression, and phase-specialized execution. Together, these designs reduce KV memory usage by while preserving efficient attention execution, delivering up to higher throughput and a reduction in TTFT, with less than 1% average accuracy loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.