acceptodds
Under review as a conference paper at ICLR 2027

Efficient LLM Inference Rescheduling via Progressive KV Cache Migration

Abstract

In multi-instance large language model (LLM) serving, dynamic request arrivals and unpredictable autoregressive decoding can create runtime resource inefficiencies such as GPU external memory fragmentation and load imbalance, complicating low-latency serving under high load. Runtime rescheduling can mitigate such inefficiencies by moving active requests across GPUs, but existing live-migration mechanisms require the full KV cache to be transferred before the source-side memory can be released. As a result, rescheduling decisions may take long to become effective, limiting their ability to promptly adjust resource allocation under evolving resource pressure. This paper revisits whether destination-side decoding must wait for full KV cache transfer. We observe temporal sparsity stability, where partial KV cache remains important over a short decoding window. With this property, the destination can resume decoding after receiving the temporally stable critical KV cache, while the remaining KV cache is refilled in the background. This enables earlier request handoff and source-side memory reclamation. So we design KVFlow, a rescheduling system for progressive KV cache migration. KVFlow exposes lightweight block-level attention statistics, selects migration-friendly requests with stable critical KV cache, and progressively refills the remaining KV cache after handoff. Our evaluation shows that, compared to a state-of-the-art rescheduling system, KVFlow reduces migration handoff latency by up to 6.89 and lowers migration-time KV cache footprint by about 30%, while preserving generation quality. In the main mixed-workload evaluation, compared with Llumnix, KVFlow reduces mean time-to-first-token (TTFT) by up to 24.0% and reduces P99 request latency and P99 TTFT by 27.1% and 47.3% on average, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.