DuetKV: Complementing Exact KV Retention with Progressive Episodic Latent Traces
Abstract
Query-agnostic key–value (KV) cache compression reduces the cache once for reuse across future queries, yet accuracy degrades sharply under tight budgets. Selective eviction preserves the original KV states but retains only a sparse token-aligned subset, while KV compaction constructs new representations at the cost of precision. Recent complementary approaches combine both, yet derive the synthetic component from the retention outcome or only after full context prefill. This leaves open how to construct complementary representations progressively from complete pre-eviction states while preserving a fixed KV budget. We introduce DuetKV, which pairs sparse exact retention with progressively constructed episodic latent traces under a fixed KV budget. For each context segment, a learned constructor synthesizes traces from its complete pre-eviction hidden states without conditioning on the current retention outcome. Frozen backbone projections map these traces into native KV states that participate alongside retained exact states in standard attention. The constructor is trained offline by self-distillation from a full-cache teacher, while the backbone and base eviction policy remain frozen. We evaluate DuetKV across 4 long-context backbones, 3 base eviction policies, and 12 prefill-intensive tasks. Across policy-budget settings, DuetKV improves the task-averaged score in 17 of 18 cases and recovers Full-KV performance with Fast KVzip at 5X compression. On 170K-token Retr.KV with Qwen2.5-14B-Instruct-1M, it raises Fast KVzip from 31.4 to 80.2 at the same 5.7 GiB peak KV memory with only 1.7% additional prefill latency. The same constructor also improves decoding-intensive reasoning under continuous KV compression without decoding-specific retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.