Recent KV Cache Optimization Strategies for Efficient Large Language Model Inference
Abstract
Autoregressive inference in Large Language Models (LLMs) is bottle necked by the memory capacity and bandwidth demands of the Key-Value (KV) cache. Although many KV cache optimization strategies have been proposed, making optimal deployment choices is difficult due to inconsistent evaluation practices and difficult to make direct comparisons. We present a systematic review and quantitative meta analysis of KV cache optimization strategies. We organized these approaches families of strategies, including quantization, token eviction, architectural reductions, and paged memory management. By extracting and normalizing the reported data from existing studies, we examined the impact of memory footprint, latency, throughput, and model quality on KV cache compression strategies. Our findings show that there is no single method that is universally superior; rather, the optimal strategy selection is highly dependent on the specific deployment constraints. We provide an evidence based recommendation matrix to guide future deployments and highlight gaps in current evaluation methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.