acceptodds
Under review as a conference paper at ICLR 2027

LatchQ: Latching onto Hot-Cache Query Interactions for KV Cache Quantization

Abstract

KV cache quantization reduces the memory footprint of large language model (LLM) inference, but minimizing reconstruction error alone may fail to preserve attention behavior. Although future queries are unavailable at quantization time, keys can interact with subsequent queries while residing in the full-precision hot cache. Motivated by this observation, we introduce ***LatchQ***, which uses these observed interactions to guide key clipping while preserving the baseline bit allocation. *QTrace* accumulates attention-weighted query energy for each key and channel, preserving differences in observed key usage and channel sensitivity. *QClip* leverages this evidence to select shared clipping ranges by minimizing interaction-weighted reconstruction error. Our method preserves the baseline storage format and cache schedule without extending hot-cache residency. Our analysis reveals 63% agreement between clipping range selections based on hot-cache observations and those based on both observed and future queries. Furthermore, ***LatchQ*** achieves the highest average LongBench scores among evaluated KV quantizers at 3.0 and 2.25 bits per dimension (bpd). It also outperforms the baseline quantizer on reasoning and code generation. Code is included in the supplementary material, and the project page is available at https://anonymous-latchq.github.io/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.