LatchQ: Latching onto Hot-Cache Query Interactions for KV Cache Quantization
Abstract
KV cache quantization reduces the memory footprint of large language model (LLM) inference, but minimizing reconstruction error alone may fail to preserve attention behavior. Although future queries are unavailable at quantization time, keys can interact with subsequent queries while residing in the full-precision hot cache. Motivated by this observation, we introduce ***LatchQ***, which uses these observed interactions to guide key clipping while preserving the baseline bit allocation. *QTrace* accumulates attention-weighted query energy for each key and channel, preserving differences in observed key usage and channel sensitivity. *QClip* leverages this evidence to select shared clipping ranges by minimizing interaction-weighted reconstruction error. Our method preserves the baseline storage format and cache schedule without extending hot-cache residency. Our analysis reveals 63% agreement between clipping range selections based on hot-cache observations and those based on both observed and future queries. Furthermore, ***LatchQ*** achieves the highest average LongBench scores among evaluated KV quantizers at 3.0 and 2.25 bits per dimension (bpd). It also outperforms the baseline quantizer on reasoning and code generation. Code is included in the supplementary material, and the project page is available at https://anonymous-latchq.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.