KV Cache Eviction as System One Memory Control for System Two Reasoning
Abstract
Long-chain reasoning expands the key–value (KV) cache, forcing memory-allocation decisions before the future roles of intermediate reasoning states are clear. We frame KV cache eviction as System One memory control for System Two reasoning: satisfying a fixed memory budget while avoiding premature semantic commitments. We propose a framework that uses relations within the evolving reasoning trajectory to selectively guide retention, rather than assigning a fine-grained importance score to every cached entry. It identifies active dependencies and explicitly superseded reasoning, while withholding semantic preferences when their roles remain unresolved. These judgments serve as revisable constraints on budget-driven eviction rather than direct deletion commands, with a simple stochastic policy resolving the remaining allocation decisions. Semantic audit runs asynchronously with generation, keeping the eviction path lightweight. Experiments on mathematical reasoning benchmarks across multiple reasoning models show higher mean accuracy than score-based and prompt-protected random eviction at matched cache budgets. An efficiency evaluation further shows reduced cache memory relative to full attention, with little additional eviction-path overhead over random eviction. These results suggest that selective, revisable semantic guidance can improve memory allocation for extended reasoning without requiring exhaustive importance estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.