acceptodds
Under review as a conference paper at ICLR 2027

PARSEE: Paging-Aware KV Cache Compression with Shared Semantic Support

Abstract

Long-context language-model serving requires large key-value (KV) caches, consuming GPU memory and limiting request concurrency. Although token eviction reduces logical cache size, heterogeneous per-head demands and page granularity can prevent these reductions from translating into reusable physical capacity. To address this challenge, we propose PARSEE, a paging-aware KV cache compression framework with shared semantic support. PARSEE is based on the observation that unequal retained counts among members sharing a page group inflate storage costs, and that shared semantic scores can coordinate these counts while preserving independent token identities. It selects shared semantic and head-specific question evidence through disjoint quotas, then allocates indivisible pages by marginal utility under a physical budget. Furthermore, a paged execution backend packs selected entries, transfers page ownership, and reclaims dense blocks, enabling batched attention over request–head groups with different retained lengths. Extensive experiments on five representative models from the Qwen and Llama families, covering dense, grouped-query, hybrid, and mixture-of-experts architectures, show that PARSEE retains near-lossless SCBench performance on Qwen3.8-27B while compressing 90% of prefill-context KV. Its Qwen3-8B backend achieves up to the throughput of uncompressed vLLM inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.