acceptodds
Under review as a conference paper at ICLR 2027

CARRIER to CARRY: Modeling KV Cache Dynamics for Provably Stable Routing in Distributed LLM Inference

Abstract

In distributed large language model (LLM) serving, key-value (KV) cache reuse reduces time-to-first-token (TTFT) by avoiding redundant prefill computation, yet exploiting shared prefixes remains challenging because caching affinity tends to concentrate requests on a few instances while pending requests alter the cached context available for subsequent reuse. To formalize this challenge, we introduce CARRIER, a mathematical model to represent request arrivals by affinity group and distributed LLM serving system with cache reuse, retention, and eviction across instances. Building on CARRIER, we develop CARRY, a routing algorithm that learns serving potential, a state function that balances immediate prefix reuse against instance load while accounting for future cache usefulness, thereby enabling a unified theoretical analysis that establishes exponential system stability, proves robustness to bounded perturbations, and quantifies CARRY's capacity advantage over baselines. Experiments across three benchmarks show that CARRY achieves improvements of up to 75.2% cache hit rate, 82.1% and 68.6% of TTFT reductions at P50 and P90, 77.8% and 68.6% of latency reductions at P50 and P90 over the state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.