CarveKV: Cross-Layer Adaptive Latent Vector Eviction for Efficient LLM Inference
Abstract
Multi-Head Latent Attention (MLA) reduces the per-token memory footprint of the KV cache, yet cache storage still grows linearly with context length. Existing token-eviction methods are largely formulated for expanded KV representations or make independent layer-wise decisions, leaving open whether eviction can be performed directly and consistently in MLA's compressed latent space. We show that latent-token importance rankings persist substantially across Transformer depth, even after controlling for causal eligibility and token position. Based on this observation, we propose CarveKV, a training-free MLA-native eviction framework that aggregates rank-normalized importance estimates across multiple layers and combines prompt-wide and final-window attention views through effective-support routing. CarveKV produces a single shared keep set across layers and scores compressed latent and decoupled-RoPE caches directly, without persistently reconstructing expanded per-head KV tensors. Across 80 natural-language and code windows, eligibility-normalized pairwise Spearman correlation reaches 0.651-0.654 and remains above 0.50 under within-position controls. On eight LongBench tasks at a matched 60% physical-cache budget, CarveKV retains 96.6% of full-cache quality. Under matched 8K-32K system settings, it also reduces prefill latency by 1.0-3.0%, with comparable decode throughput and peak memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.