Fewer Bits, More Tokens: Rethinking KV Cache Quantization in Agentic Coding
Abstract
Agentic coding has emerged as a prominent application of LLMs, but its multi-turn interaction trajectories continually expand the KV cache, imposing substantial memory demands during inference. KV cache quantization reduces per-token memory usage, but do these savings translate into more efficient agent trajectories? We find that low-bit quantization disproportionately amplifies three failures: repetitive action loops, improper tool use, and insufficient debugging, while increasing the actions and tokens required per task. To identify which attention heads are most sensitive to quantization around these failures, we measure changes in attention distributions and head outputs under controlled KV cache perturbations. Highly sensitive heads comprise only about 5% of all heads and cluster in a few layers. Guided by this localization, we introduce a simple and effective mixed-precision strategy that restores higher precision in selected layers while retaining low precision elsewhere. On SWE-bench Verified with Qwen3-Coder-30B-A3B-Instruct, increasing average precision from to bits substantially reduces the three behavioral failures and brings task resolution close to BF16. Despite higher per-token precision, observed KV cache usage per case decreases by 12% versus uniform 3-bit quantization. These results shift the focus of agent quantization from storing each token more cheaply to completing tasks more reliably with less memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.