PARSEE: Paging-Aware KV Cache Compression with Shared Semantic Support
Abstract
Long-context language-model serving requires large key-value (KV) caches, consuming GPU memory and limiting request concurrency. Although token eviction reduces logical cache size, heterogeneous per-head demands and page granularity can prevent these reductions from translating into reusable physical capacity. To address this challenge, we propose PARSEE, a paging-aware KV cache compression framework with shared semantic support. PARSEE is based on the observation that unequal retained counts among members sharing a page group inflate storage costs, and that shared semantic scores can coordinate these counts while preserving independent token identities. It selects shared semantic and head-specific question evidence through disjoint quotas, then allocates indivisible pages by marginal utility under a physical budget. Furthermore, a paged execution backend packs selected entries, transfers page ownership, and reclaims dense blocks, enabling batched attention over request–head groups with different retained lengths. Extensive experiments on five representative models from the Qwen and Llama families, covering dense, grouped-query, hybrid, and mixture-of-experts architectures, show that PARSEE retains near-lossless SCBench performance on Qwen3.8-27B while compressing 90% of prefill-context KV. Its Qwen3-8B backend achieves up to the throughput of uncompressed vLLM inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.