acceptodds
Under review as a conference paper at ICLR 2027

QEvict: Retain Broadly, Read Selectively for Long-Context Decoding

Abstract

Long-context LLM inference is increasingly constrained by the memory and computation of the key-value (KV) cache. We find that KV utility exhibits two distinct timescales: long-horizon historical importance is broad and evolves gradually, while relevance to an individual query is much more concentrated. Motivated by this, we propose QEvict, a window-based KV-cache architecture that separates historical residency, storage representation, and current-query access. Historical windows are periodically routed under a byte budget among execution-datatype storage, compact INT2 storage, and permanent removal using accumulated attention. A write-once INT2 representation supports stable migration across storage priorities, while a compact query-conditioned representation selects only a fraction of retained low-bit windows for token-level attention. Unread windows remain represented through calibrated window-level completion. A fused mixed-precision kernel realizes selective INT2 access without materializing the full compressed cache. Across LongBench, RULER, and GSM8K with Llama and Mistral, QEvict improves the quality–memory trade-off over representative eviction, quantization, and mixed-precision baselines, with particularly strong gains under tight KV-cache budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.