acceptodds
Under review as a conference paper at ICLR 2027

PagePrune: Bridging the Pruning–Execution Gap in Paged-Attention KV Cache

Abstract

Long-context large language model (LLM) inference is increasingly constrained by the growing Key-Value (KV) cache, which consumes substantial GPU memory and incurs repeated memory traffic during decoding. Existing KV-cache pruning methods reduce the retained cache size, but often overlook how the surviving KV entries are organized and accessed by attention kernels, limiting the resulting end-to-end (E2E) gains. We present **PagePrune**, an execution-aware, page-granular KV-cache pruning system for vLLM that jointly optimizes KV selection and execution layout. PagePrune periodically ranks KV pages using exponentially smoothed attention importance and retains high-utility pages under a configurable budget, while protecting critical prefix and recent context. It further aligns pruning decisions with attention-kernel tiles and performs compact-once repacking to expose the retained cache as a contiguous working set. We integrate PagePrune into vLLM and evaluate it across long-context language, vision question answering, and video-language workloads. Empirically, PagePrune substantially reduces KV-cache memory footprint while achieving up to **7.91×** E2E speedup over vanilla vLLM with comparable task quality. Experiments and analysis demonstrate that effective KV-cache pruning requires not only selecting what to retain, but also organizing how the retained cache is executed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.