KVLease: Dynamic KV Memory Sharing for Mixture-of-Models Inference
Abstract
Language-model applications can combine several models to answer one request. In this Mixture of Models (MoM) setting, workflows connect parallel and sequential model calls through their intermediate outputs. Stage transitions shift key–value (KV) cache demand, but overlapping requests keep earlier-stage engines busy. A downstream model can therefore have its inputs ready yet wait for capacity that those engines repeatedly reuse, delaying the complete answer. Fixed partitions strand capacity; lending idle memory alone cannot recover pages from busy borrowers. The challenge is to serve the next stage without interrupting existing uses. We present KVLease, a KV-memory sharing system designed for MoM workflows. It couples lending and per-model allocation guarantees with targeted return: selected borrowed pages stop accepting new allocations, then return after existing references and GPU accesses finish. Capacity is shared per GPU; model-specific KV contents stay private. Our analysis characterizes how workflow stages correlate memory demand, separates partition imbalance from temporal sharing, and establishes conditional return bounds. Replicated evaluations across heterogeneous GPU platforms, serving engines (e.g., vLLM and SGLang), and diverse MoM workloads show 28–53% lower end-to-end p95 latency than equal-token fixed allocations at the reported memory-constrained operating points, over 50% higher completed-request throughput, and up to 35 percentage points higher service-level objective (SLO) attainment. Against kvcached at matched capacity, p95 falls a further 24–31% at high load. These gains preserve model computation, with identical outputs in paired correctness checks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.