acceptodds
Under review as a conference paper at ICLR 2027

TridentKV: Optimizing KV Cache Eviction with Diversity-Aware Budget Allocation and Spatial-Temporal Score Aggregation

Abstract

Long-context inference in large language models (LLMs) is increasingly bottlenecked by the growing memory and computation cost of the key-value (KV) cache. While KV cache eviction provides a training-free solution, existing methods typically optimize cache allocation or importance aggregation without explicitly accounting for behavioral diversity across layers and attention heads or localized variations in token importance. We introduce TridentKV, a training-free KV cache eviction framework that improves cache utilization through diversity-aware budget allocation and spatial-temporal score aggregation. TridentKV consists of three complementary components: (1) uniqueness-based layer budget allocation, which assigns larger cache budgets to layers exhibiting distinctive attention behaviors; (2) behavior-guided head budget allocation, which clusters attention heads by behavioral similarity and restricts budget competition within each cluster to preserve diverse attention patterns; and (3) spatial-temporal token score aggregation, which jointly captures importance variations across head clusters and query window to preserve locally salient tokens. Experiments on LongBench, RULER, and Needle-in-a-Haystack across Llama, Mistral, and Qwen models show that TridentKV robustly preserves long-context performance under constrained cache budgets. With a cache budget of 512, it retains over 99.6% of full-cache performance on LongBench and outperforms the strongest evaluated baselines on 32K RULER by 2.7%–5.4%, while achieving up to 2.76 decoding speedup and a 36.2% reduction in peak memory at 128K context length. We provide our code in the supplementary material to facilitate straightforward reproduction of our method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.