Safety in Numbers: The KV Cache Was Never That Fragile to Repeated Pruning
Abstract
Long-running agents accumulate stale context from failed retrievals and superseded reasoning, occupying GPU memory and distorting attention. Removing it at the text layer forces a re-encode; at the KV-cache layer it is nearly free: excise rows, re-rotate survivor positions, continue. Whether that composite cache behaves like an honest re-prefill of the surviving tokens is untested, and production practice stays conservative. We test it directly, holding surviving content byte-identical and varying only operation count (Q1), window position (Q2), and eviction timing (Q3); relevance is oracle-given and the mechanism a reused RoPE rotation, so the operation itself is the only remaining variable. The caution inverts: repeating the cheap operation makes the result more faithful, not less. Over 8 rounds at 92.9% pruning, composite-KV quality reaches 91.3% of a clean re-prefill and improves with round count, with no drift over 26+ real-trajectory rounds (Q1). Re-indexing costs only 0.8% and matters only when survivor positions exceed the model's trained window, confirmed by a pre-registered model swap (Q2). Timing is the genuine open variable (Q3): correctness slips only a little as junk lingers, a small, task-dependent cost that timely pruning keeps low. The worry that repeated pruning compounds error is unsupported; it is timing, not repetition, that costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.