acceptodds
Under review as a conference paper at ICLR 2027

Position-Aware Cooperative Token-Channel Compression with Key-Value Decoupling

Abstract

The linearly growing key-value (KV) cache imposes a critical storage bottleneck for large language models (LLMs) in long-context inference scenarios, hindering their practical deployment. Existing methods compress the KV cache by evicting or merging tokens and reducing channel redundancy, thereby reducing its memory footprint. However, these merging methods overlook that after Rotary Position Embedding (RoPE), semantically similar tokens can yield inconsistent attention responses due to positional differences, whereas tokens with similar attention responses have higher positional-function similarity and are more mergeable. Moreover, these methods handle the token and channel dimensions in isolation, and apply a uniform compression strategy to both K and V, overlooking token-channel cooperative potential and distinct roles of K and V in the attention mechanism. In this paper, we propose a Position-Aware Cooperative Token-Channel Compression with Key-Value Decoupling method (PACT-KV). It designs positional-function similarity to measure the mergeability of tokens within the SimHash cluster for K-side compression, which jointly considers semantics and attentions, and avoids positional interference brought by physical proximity. Meanwhile we propose a cluster-conditioned K-side channel pruning to capture consistent channel-energy patterns in the same cluster, which naturally cooperates with the K-side token merging to reduce the overhead. Due to the semantic aggregation role of V, which contrasts with the query-matching role of K, PACT-KV further performs an independent V-side token merging and executes V-side channel pruning when V tokens still exceed the budget. Experimental results on LongBench and Needle-in-a-Haystack (NIAH) show that PACT-KV achieves strong performance while significantly reducing both memory footprint and inference latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.