acceptodds
Under review as a conference paper at ICLR 2027

CacheFlow: Efficient LLM Serving via Automated 3D-Parallel KV Cache Restoration

Abstract

KV cache restoration is becoming a major bottleneck in long-context LLM serving, including multi-turn conversations, agentic workloads, and RAG. Existing advances largely treat restoration as a coarse per-request choice between recomputing cached states and loading them from external storage (e.g., CPU memory or remote machines), overlooking structured parallelism within the model and resource contention across batched requests. Increasingly, modern hybrid architectures complicate restoration as their interleaved attention and recurrent layers introduce heterogeneous compute costs and state footprints. We present CacheFlow, which rethinks restoration as a multi-dimensional parallel execution problem. To our knowledge, CacheFlow is the first framework to jointly exploit restoration parallelism across tokens, layers, and GPUs. CacheFlow unifies token- and layer-wise restoration as a staircase partition of the token–layer space, while lightweight boundary activations enable concurrent restoration across model shards. A batch-aware dynamic-programming planner jointly determines what to recompute and what to load across requests and layer blocks, balancing GPU computation against shared I/O resources. Across dense and hybrid models spanning single-digit to hundred-billion parameters, diverse workloads, and hardware conditions, CacheFlow reduces restoration latency by 2.24–3.00 on average over existing advances, directly improving end-to-end serving latency by 1.64.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.