CacheFlow: Efficient LLM Serving via Automated 3D-Parallel KV Cache Restoration
Abstract
KV cache restoration is becoming a major bottleneck in long-context LLM serving, including multi-turn conversations, agentic workloads, and RAG. Existing advances largely treat restoration as a coarse per-request choice between recomputing cached states and loading them from external storage (e.g., CPU memory or remote machines), overlooking structured parallelism within the model and resource contention across batched requests. Increasingly, modern hybrid architectures complicate restoration as their interleaved attention and recurrent layers introduce heterogeneous compute costs and state footprints. We present CacheFlow, which rethinks restoration as a multi-dimensional parallel execution problem. To our knowledge, CacheFlow is the first framework to jointly exploit restoration parallelism across tokens, layers, and GPUs. CacheFlow unifies token- and layer-wise restoration as a staircase partition of the token–layer space, while lightweight boundary activations enable concurrent restoration across model shards. A batch-aware dynamic-programming planner jointly determines what to recompute and what to load across requests and layer blocks, balancing GPU computation against shared I/O resources. Across dense and hybrid models spanning single-digit to hundred-billion parameters, diverse workloads, and hardware conditions, CacheFlow reduces restoration latency by 2.24–3.00 on average over existing advances, directly improving end-to-end serving latency by 1.64.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.