HarmonyTransfer: KVCache Transfer Scheduling for Efficient LLM Agent Serving
Abstract
LLM agent applications often contain fan-out stages whose completion depends on multiple requests. Such a stage reaches collective prefill completion only when every required request has returned its first token. Serving long-context fan-out stages efficiently requires moving the reusable KVCache from lower-tier storage back onto the GPU. That transfer runs over a link shared by every live stage. Existing systems, however, schedule KVCache transfers independently for each request, without accounting for the stage's joint completion semantics. We introduce the Agent Prefill Group (APG), a first-class abstraction representing the set of requests whose first-token completion jointly determines a stage's collective prefill completion, and present HarmonyTransfer, an APG-aware KVCache transfer scheduling system. HarmonyTransfer implements objective-specific APG scheduling policies for reducing mean completion time and improving deadline attainment. It dispatches KVCache transfers in bounded chunk-level batches, allowing the scheduler to make a new APG scheduling decision at each batch boundary rather than only at full-request transfer completion. Built on LMCache and vLLM, HarmonyTransfer reduces mean APG completion time by up to 51.6% over the best-performing baseline and improves APG deadline attainment by up to 23.0 percentage points in testbed experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.