QEvict: Retain Broadly, Read Selectively for Long-Context Decoding
Abstract
Long-context LLM inference is increasingly constrained by the memory and computation of the key-value (KV) cache. We find that KV utility exhibits two distinct timescales: long-horizon historical importance is broad and evolves gradually, while relevance to an individual query is much more concentrated. Motivated by this, we propose QEvict, a window-based KV-cache architecture that separates historical residency, storage representation, and current-query access. Historical windows are periodically routed under a byte budget among execution-datatype storage, compact INT2 storage, and permanent removal using accumulated attention. A write-once INT2 representation supports stable migration across storage priorities, while a compact query-conditioned representation selects only a fraction of retained low-bit windows for token-level attention. Unread windows remain represented through calibrated window-level completion. A fused mixed-precision kernel realizes selective INT2 access without materializing the full compressed cache. Across LongBench, RULER, and GSM8K with Llama and Mistral, QEvict improves the quality–memory trade-off over representative eviction, quantization, and mixed-precision baselines, with particularly strong gains under tight KV-cache budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.