acceptodds
Under review as a conference paper at ICLR 2027

CacheFlow: Fast and Memory Efficient LLM Inference via In-Network Prefix Cache Pool

Abstract

Agentic systems consist of multiple copies of the same agent, defined by a common prompt prefix, that are hosted on multiple physical machines. Each of the machines is required to process the same common prefixes separately, leading to redundant computation across and increased latency. We recognize that all of the physical machines hosting the agents are connected by a computer network, making the network the trivial choice when storing the shared data needed across all machines. Leveraging this idea, this paper presents CacheFlow, a design and an implementation of a prefix cache pool that stores KV cache on a network device and delivers it to the GPU via GPUDirect RDMA bypassing the host CPU and DRAM. Our evaluations show that CacheFlow improves the overall prefix cache efficiency compared to vanilla vLLM by reducing the average Time-to-First-Token (TTFT) by up to 82%, increasing throughput by 39%, reducing the host memory storage requirement by 50%. In addition, CacheFlow shows up to 3X improvement in I/O throughput compared to host DRAM based cache systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.