acceptodds
Under review as a conference paper at ICLR 2027

From Memory Savings to Throughput Gains: Serving-Native KV Compression for Sparse-Attention LLMs

Abstract

Long-context LLM serving is increasingly bottlenecked by the growing memory and bandwidth demands of the KV cache. Sparse attention alleviates the KV access cost, but still leaves the full KV cache resident, limiting serving concurrency. Existing KV compression methods primarily target compact storage or optimized kernels, rather than the execution and lifecycle of compressed KV in online serving, where compression timing, efficient attention over compressed states, and cache management directly affect performance. We present Duet, a serving-native KV compression framework that turns KV memory savings into practical serving gains while preserving the underlying sparse-attention selection semantics. Duet first develops an inference-friendly low-rank representation that exploits cross-head redundancy and removes per-token reconstruction from sparse attention. Building on this, a mixed native-latent attention operator directly consumes selected KV blocks in either form without materializing compressed history. Finally, Duet integrates compression into the serving runtime through dual-representation KV management, watermark-triggered batched compression, and phase-specialized execution. Together, these designs reduce KV memory usage by while preserving efficient attention execution, delivering up to higher throughput and a reduction in TTFT, with less than 1% average accuracy loss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.