PagePrune: Bridging the Pruning–Execution Gap in Paged-Attention KV Cache
Abstract
Long-context large language model (LLM) inference is increasingly constrained by the growing Key-Value (KV) cache, which consumes substantial GPU memory and incurs repeated memory traffic during decoding. Existing KV-cache pruning methods reduce the retained cache size, but often overlook how the surviving KV entries are organized and accessed by attention kernels, limiting the resulting end-to-end (E2E) gains. We present **PagePrune**, an execution-aware, page-granular KV-cache pruning system for vLLM that jointly optimizes KV selection and execution layout. PagePrune periodically ranks KV pages using exponentially smoothed attention importance and retains high-utility pages under a configurable budget, while protecting critical prefix and recent context. It further aligns pruning decisions with attention-kernel tiles and performs compact-once repacking to expose the retained cache as a contiguous working set. We integrate PagePrune into vLLM and evaluate it across long-context language, vision question answering, and video-language workloads. Empirically, PagePrune substantially reduces KV-cache memory footprint while achieving up to **7.91×** E2E speedup over vanilla vLLM with comparable task quality. Experiments and analysis demonstrate that effective KV-cache pruning requires not only selecting what to retain, but also organizing how the retained cache is executed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.