acceptodds
Under review as a conference paper at ICLR 2027

LAB-Sparse: Local Accumulative Boost for Adaptive Semantic-Block Retention in KV Cache

Abstract

The expanding context windows of large language models (LLMs) have brought substantial gains in reasoning and comprehension, yet this expansion comes at the cost of a linearly growing Key-Value (KV) cache that strains GPU memory and increases decoding latency.Existing KV sparsification methods, whether heuristic token-level eviction or fixed-size block retention, share a common limitation: they treat the KV cache as a collection of isolated units. This overlooks a fundamental property of natural language—semantic information is typically organized as continuous spans rather than isolated tokens. Fixed-size blocks are particularly problematic: a coherent semantic unit may be split across multiple blocks, and averaging scores within a block can dilute or filter out critical tokens. Consequently, these methods often suffer from fragmented critical information and degraded continuity in multi-hop reasoning and long-context retrieval. To address this, we propose LAB-Sparse, a dynamic continuous block retention method powered by Local Accumulative Boost (LAB). Our core insight is to elevate the basic unit of sparsification from individual tokens to dynamic-length contiguous blocks, where block lengths are determined by contiguous runs above a quantile-based threshold, which implicitly reflect local importance density. The LAB mechanism applies a neighborhood cumulative gain to smooth token-level importance scores, thereby boosting tokens surrounded by high-importance neighbors and suppressing isolated noise. On this enhanced distribution, a quantile-based interval delineation strategy identifies high-value contiguous regions, followed by a greedy block-wise retention policy that prioritizes complete contiguous blocks under a fixed KV budget. Experiments on Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B across LongBench and RULER show that LAB-Sparse achieves competitive or superior average accuracy compared with strong baselines such as H2O, SnapKV, StreamingLLM, ChunkKV, and PyramidKV, while remaining close to full attention in the 4K–16K range. Under a fixed KV budget of 512, it matches the decoding throughput and memory savings of block-based baselines at 32K ( and 12.6%), while improving average accuracy over token-level methods.Anonymous code will be submitted as supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.