CARRIER to CARRY: Modeling KV Cache Dynamics for Provably Stable Routing in Distributed LLM Inference
Abstract
In distributed large language model (LLM) serving, key-value (KV) cache reuse reduces time-to-first-token (TTFT) by avoiding redundant prefill computation, yet exploiting shared prefixes remains challenging because caching affinity tends to concentrate requests on a few instances while pending requests alter the cached context available for subsequent reuse. To formalize this challenge, we introduce CARRIER, a mathematical model to represent request arrivals by affinity group and distributed LLM serving system with cache reuse, retention, and eviction across instances. Building on CARRIER, we develop CARRY, a routing algorithm that learns serving potential, a state function that balances immediate prefix reuse against instance load while accounting for future cache usefulness, thereby enabling a unified theoretical analysis that establishes exponential system stability, proves robustness to bounded perturbations, and quantifies CARRY's capacity advantage over baselines. Experiments across three benchmarks show that CARRY achieves improvements of up to 75.2% cache hit rate, 82.1% and 68.6% of TTFT reductions at P50 and P90, 77.8% and 68.6% of latency reductions at P50 and P90 over the state-of-the-art methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.