acceptodds
Under review as a conference paper at ICLR 2027

Multiple Stages, More Gains: Proactive Staged Compression for Efficient LLM Inference

Abstract

The ever-growing Key-Value (KV) Cache generated during decoding becomes a bottleneck for large language models (LLMs) in long-context processing. Recent research has mainly focused on compressing the KV Cache through eviction and merging strategies, thereby enhancing long-context inference capability. However, most existing methods rely on aggressive compression at the memory limit, resulting in cumulative attention perturbation and model performance degradation. In this paper, we propose **PSC** (**P**roactive **S**taged **C**ompression), a compressor-agnostic scheduling framework that distributes KV Cache compression across multiple stages before the cache reaches its capacity limit. Our key insight is that decomposing compression into multiple stages with small removal ratios significantly reduces attention perturbation. From the cache capacity, prompt length and target retention ratio, we analytically derive the number of compression stages and the retained KV Cache size at each stage to obtain a closed-form two-phase compression schedule. Experimental results across various models and datasets show that incorporating PSC significantly enhances the performance of existing representative compression methods, achieving consistent accuracy improvement and accelerating inference by up to 1.17 compared to representative baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.