acceptodds
Under review as a conference paper at ICLR 2027

Fewer Bits, More Tokens: Rethinking KV Cache Quantization in Agentic Coding

Abstract

Agentic coding has emerged as a prominent application of LLMs, but its multi-turn interaction trajectories continually expand the KV cache, imposing substantial memory demands during inference. KV cache quantization reduces per-token memory usage, but do these savings translate into more efficient agent trajectories? We find that low-bit quantization disproportionately amplifies three failures: repetitive action loops, improper tool use, and insufficient debugging, while increasing the actions and tokens required per task. To identify which attention heads are most sensitive to quantization around these failures, we measure changes in attention distributions and head outputs under controlled KV cache perturbations. Highly sensitive heads comprise only about 5% of all heads and cluster in a few layers. Guided by this localization, we introduce a simple and effective mixed-precision strategy that restores higher precision in selected layers while retaining low precision elsewhere. On SWE-bench Verified with Qwen3-Coder-30B-A3B-Instruct, increasing average precision from to bits substantially reduces the three behavioral failures and brings task resolution close to BF16. Despite higher per-token precision, observed KV cache usage per case decreases by 12% versus uniform 3-bit quantization. These results shift the focus of agent quantization from storing each token more cheaply to completing tasks more reliably with less memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.