acceptodds
Under review as a conference paper at ICLR 2027

CheeSe : Channel Permutation for Semi-structured KV Cache Pruning

Abstract

The KV cache is the memory-bandwidth bottleneck in long-context LLM inference. Element-wise unstructured pruning preserves quality best but obtains no compute acceleration, and its storage overhead keeps the effective compression ratio below the nominal one; structured methods accelerate computation but lose accuracy quickly as compression grows. 2:4 pruning offers both acceleration on sparse tensor cores and a high effective compression ratio, yet prior work rejected it for the KV cache because of its quality loss, and subsequent work accelerated it with dedicated kernels while leaving that loss in place. We propose CheeSe, which recovers this loss with two techniques. First, we show that the loss arises almost entirely on the key cache and stems from how channels are grouped rather than from which elements are removed: removing the same number of key elements costs under unstructured pruning but under 2:4, because important channels cluster within the same group. A sorting-based channel permutation, computed once during calibration and transferable across tasks, recovers most of this loss, whereas more elaborate permutation searches from the weight pruning literature cost up to four orders of magnitude more compute and bring no further gain. Second, output-aware selection, which is harmful for 2:4 alone, becomes beneficial once the permutation is in place. Across five models, CheeSe stays within 0.21 points of unstructured pruning and recovers more of the 2:4 loss than keeping sink tokens and the recent context dense, at no cost in compression. It improves the effective compression ratio by 9.2 percentage points and accelerates decode attention by with an existing 2:4 kernel at a runtime overhead of 2.5%, and it composes with token eviction and quantization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.