acceptodds
Under review as a conference paper at ICLR 2027

Rethinking KV Cache Offloading: Bandwidth Scaling and Decode Continuity

Abstract

KV cache offloading trades data movement for memory capacity in language model serving. As offload bandwidth increases, should systems continue to prioritize minimizing transferred bytes? We study this question using a storage-only NVLink peer as a controlled surrogate for a high-bandwidth remote tier, with inference resources held fixed. A bandwidth sweep reveals a reversal: request-granular migration begins slower than two fine-grained policies but becomes faster as transfer cost falls. Load sweeps and runtime analysis connect this reversal to a broader design concern: preserving useful decoding while recovering memory capacity. We formulate three design rules and instantiate them in CargpKV: amortize migration interruptions, recover headroom with few logical actions, and bound migration activity. Across three dense models and six workload regimes, comparisons with LayerKV, NEO, ZeRO-Inference, and FlexGen reveal both gains and boundaries. At the highest tested static latency loads, CargpKV reduces P99 time per output token by 24.3–49.3% relative to SOTA. The resulting guideline is conditional and portable: a fast offload path should be designed around the exposed cost of interrupting decoding, alongside the bytes it moves.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.