acceptodds
Under review as a conference paper at ICLR 2027

KVReclaim: Recompute-Aware Placement and Migration-Free Reclamation for KV Caches on ZNS SSDs

Abstract

Large language model (LLM) inference increasingly relies on long contexts and persistent conversation histories, causing KV caches to grow rapidly and become a major bottleneck for scalable serving. SSD-backed KV caching provides a cost-effective means of extending scarce GPU and host memory, but conventional SSD accesses incur substantial KV recovery latency, especially as device-managed garbage collection (GC) induces data migration that can interfere with foreground KV I/O. Zoned Namespace (ZNS) SSDs provide an opportunity to eliminate such unnecessary migration by exposing data placement and space reclamation to the host. Based on this insight, we present KVReclaim, a host–SSD coordinated KV-cache management framework for ZNS SSDs that exploits KV semantics across the storage hierarchy. KVReclaim manages KV states at conversation granularity and flushes them into zones according to their runtime access behavior and recomputation overhead. During garbage collection (GC), it directly discards selected KV states instead of migrating them, thereby reducing GC-induced data movement. KVReclaim further dynamically selects between SSD-backed KV recovery and recomputation by comparing their estimated costs, thereby reducing recovery overhead under varying SSD I/O conditions. Experiments on long-context and multi-turn LLM workloads show that KVReclaim reduces E2E latency by 74.2% on average compared with state-of-the-art SSD-backed KV-cache management schemes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.