DynamicKV: Predictive and Reversible KV-Cache Precision Allocation for Agentic LLM Inference
Abstract
Agent workflows and long-chain reasoning continually retain and access growing histories, making the KV cache a major GPU memory cost during inference. Existing KV quantization methods typically determine precision when historical KV states enter the cache or reduce bit widths unidirectionally during generation. Analysis of real generation traces shows that earlier KV states can be reused after a period of limited attention; moreover, the effect of demoting the same KV region on subsequent inference can increase or decrease across generation stages. To anticipate changes in future demand and adapt quantization under a limited budget, we introduce DynamicKV, a predictive and reversible KV precision allocation method. DynamicKV partitions historical KV states into regions aligned with underlying storage blocks and periodically estimates two signals: , a score for significant next-interval attention, and , a score for significant damage from demotion to low-bit precision. Their product, , serves as a gated ranking score for dynamic precision allocation under a fixed proportional GPU KV-byte budget. Each region retains a low-bit base on the GPU and a separable residual in host memory. When the region becomes important again, activating the residual restores an approximately high-bit representation, enabling repeated demotion and restoration. On agent workflow tasks (-bench and BFCL V4), DynamicKV improves success rates by 5.0–10.5 percentage points over the strongest compression baseline for each model. On AIME 2024/2025 and MATH-500, it improves accuracy by up to 3.3 percentage points over the strongest compression baselines. On a single A100-80GB GPU, DynamicKV supports the concurrency of BF16 at 32K and improves aggregate output throughput by 4.7–44.3% at sequence lengths of 16K–32K.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.