KVFlow: Cross-Instance KV Cache Flow for Efficient and Balanced LLM Serving
Abstract
Multi-instance large language model (LLM) serving must balance KV cache reuse and load balancing. Prior work addresses this trade-off through request scheduling. However, scheduling cannot eliminate the fundamental conflict because hot KV cache couples reuse benefits with instance load, preventing requests from simultaneously achieving low computation time through reuse and low queueing delay through load balancing. Our key insight is to manage KV cache directly by replicating hot KV blocks across instances, preventing hot cache from causing load imbalance. We first define KV cache access pressure to characterize how cache distribution affects future request load across instances and show that balancing access pressure decouples KV cache reuse benefits from instance load. Based on this analysis, we propose KVFlow, the first KV cache management system that balances access pressure across instances. Across multiple LLMs and real LLM serving workloads, KVFlow reduces average time-to-first-token (TTFT) by 12.5%–54.3% compared with state-of-the-art production systems, including vLLM, Dynamo, and LMetric.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.