acceptodds
Under review as a conference paper at ICLR 2027

GRIK: Grouped Relevance-Integrated Key Pruning for Grouped-Query Attention LLMs

Abstract

The key–value (KV) cache is a major memory bottleneck in long-context inference, and existing K-channel pruning does not shrink it under grouped-query attention (GQA): in the evaluated implementations, ThinK materializes K and V per query head ( total KV) and SparK reconstructs a full-width cache (). We identify two structural conditions, a mask shared across the query heads of each KV group and across all cached tokens, under which channel pruning yields a dense shared-key cache whose width equals the retained channel count. We introduce GRIK, a training- and calibration-free method that satisfies both conditions by construction and derives a prompt-adaptive mask per KV head from output-projection influence, key-projection norms, and attention energy. At 40% K-channel pruning, GRIK stores K at and total K+V at the bf16 baseline. Without token eviction, its LongBench average is within points of Full KV on three GQA backbones; with H2O eviction it is the best non-Full-KV method in 19 of 24 (model, budget, ratio) cells, and at approximately matched total-KV memory it improves over H2O alone by – points. The reduction composes with other axes: adding GRIK to SnapKV, FP8, or int4 KV changes LongBench by at most points while cutting total KV to as low as . On RULER, its eight exact-retrieval tasks average at least through 32K, although the 13-task average trails Full KV by – points. In a fixed-memory H100 serving test at 50% pruning, the smaller cache admits more concurrent sequences and raises aggregate request throughput by , with per-user decode latency that of the baseline at the higher concurrency. Code and configurations are available at https://anonymous.4open.science/r/GRIK/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.