acceptodds
Under review as a conference paper at ICLR 2027

Recent KV Cache Optimization Strategies for Efficient Large Language Model Inference

Abstract

Autoregressive inference in Large Language Models (LLMs) is bottle necked by the memory capacity and bandwidth demands of the Key-Value (KV) cache. Although many KV cache optimization strategies have been proposed, making optimal deployment choices is difficult due to inconsistent evaluation practices and difficult to make direct comparisons. We present a systematic review and quantitative meta analysis of KV cache optimization strategies. We organized these approaches families of strategies, including quantization, token eviction, architectural reductions, and paged memory management. By extracting and normalizing the reported data from existing studies, we examined the impact of memory footprint, latency, throughput, and model quality on KV cache compression strategies. Our findings show that there is no single method that is universally superior; rather, the optimal strategy selection is highly dependent on the specific deployment constraints. We provide an evidence based recommendation matrix to guide future deployments and highlight gaps in current evaluation methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.