acceptodds
Under review as a conference paper at ICLR 2027

The Hit Is (Almost) Always Cheaper: Prefix-Cache Economics on Bandwidth-Rich, FLOPs-Poor Accelerators

Abstract

Automatic prefix caching (APC) is widely described as an almost-free win: when a request shares a prefix with a prior one, the engine skips recomputing the shared tokens’ key/value (KV) cache. We present a controlled measurement study of what a prefix-cache hit actually costs and buys on a bandwidth-rich, FLOPs-poor accelerator (NVIDIA H20-141GB), where folklore calibrated on compute-rich parts may not transfer. Across Qwen2.5 models at 0.5B–7B, contexts up to 32K, and batches 1–64, we fit a two-term cost model that separates the compute-bound recompute path from the bandwidth-bound hit path, and validate it end-to-end on vLLM. Three findings organize the results. First, hits win overwhelmingly at realistic prefix lengths — up to 169× lower suffix-prefill latency at 32K context — and no measured regime inverts (512 tokens, the grid’s shortest, already sits above the fitted crossover); the honest boundary of the win is a short-prefix crossover (locally fitted at 152–635 tokens at batch 1; a slope-poor interpolation, with the win settling near 1.7× at 128-token prefixes end-to-end) and a fixed 15–21 ms hit-path floor. Second, the hit path is not free even when it wins: materializing paged KV state with a 16-token-block gather achieves only 0.135–0.141 TB/s at small working sets (0.072 TB/s at 30 GB) versus 1.10 TB/s for a contiguous device copy — an 8–15× penalty on any back-end that materializes paged KV through gather-like scattered reads (engines that share physical blocks, such as vLLM’s APC, avoid this path). Third, 141 GB capacity declines with workload skew but does not vanish: under Zipf popularity with exponent 1.0, a 30 GB LRU budget still leaves 21 points of hit rate on the table versus 106 GB; only at exponent 1.4 does the gap close to 9 points. We release all harnesses and raw results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.