acceptodds
Under review as a conference paper at ICLR 2027

Coordinated Mixed-Precision KV Cache Quantization for Efficient LLM-based Multi-agent Systems

Abstract

LLM-based multi-agent systems solve complex tasks through problem decomposition, specialized tool use, and iterative verification. However, collaboration produces multiple growing key-value (KV) caches that can coexist in memory, making cache storage a bottleneck. KV-cache quantization reduces this footprint, but independently applying single-model policies ignores differences in agents' precision needs. This can over-compress sensitive computations while spending precision on more tolerant caches, potentially degrading collective reasoning. In this paper, we introduce **Co**ordinated **Mix**ed-Precision **KV** Cache Quantization (**CoMixKV**), a novel method that integrates online sensitivity estimation with joint bit-width allocation under a shared memory budget. To our knowledge, **CoMixKV** is the first mixed-precision KV-cache quantization framework to jointly allocate precision across simultaneously resident agent caches under such a budget. An online estimator derives each agent's sensitivity from a local attention-error bound using recent queries, keys, and values, without calibration data, trial quantization, or additional model passes. A shared controller combines these scores with remaining quantizable capacity and byte-level storage costs to minimize total estimated error. Allocations are updated as agents start or finish and fresh observations become available. Selected widths apply only to newly quantized blocks; recent states retain native precision and previously quantized blocks remain unchanged. Across four benchmarks and three LLM backbones at average 2-bit, 3-bit, and 4-bit KV-cache settings, **CoMixKV** achieves strong accuracy compared with the evaluated quantization baselines. At the average 3-bit setting on Gemma-4-31B-it, **CoMixKV** achieves 68.35% average accuracy, outperforming the strongest quantized baseline by 0.59%. On Qwen3-32B, **CoMixKV** further reduces peak KV-cache memory by 32.08% and achieves a 1.38 inference speedup over FullKV.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.