acceptodds
Under review as a conference paper at ICLR 2027

ARI: Adaptive Rule-Interpolated KV Cache Migration Scheduling for Multi-GPU LLM Serving

Abstract

Large language model serving maintains a request-level key-value (KV) cache that grows during decoding. Although PagedAttention reduces within-instance fragmentation, request growth and completion can imbalance multi-GPU placement. Existing routing and migration mechanisms leave request-level migration value under interacting memory, transfer-cost, and remaining-duration conditions insufficiently modeled. We present , an adaptive rule-interpolated decision layer for runtime Live-KV migration. After a global imbalance detector fires, candidates are represented by source pressure, KV size, estimated migration cost, global imbalance, and remaining execution time. Gaussian rule matching and normalized consequent interpolation yield continuous migration priorities. A separate destination controller plans feasible moves on a virtual cluster state, stopping at the first balanced prefix along its greedy path (minimum Top-) when one exists. The priority rule base grows only in dense, insufficiently covered regions, and observed post-migration utility updates rule confidence. retains MELL's initial allocation and reuses request-level Direct-KV or Token-RePrefill execution paths. Across six model–platform settings on 4 and 8 GPUs, achieves higher output-token throughput than MELL. On 4-GPU Llama-2-7b-chat, throughput increases from 436.82 to 675.31 tokens/s (54.6%) and P95 latency decreases by 30.9%. In selected settings, also improves serving performance or migration efficiency over other baselines. Upon acceptance, we will publicly release all source code and experimental scripts to support reproducibility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.