AdmitKV: Provenance- and Type-Gated KV Transfer for Edited Agent Contexts
Abstract
An identity-cache miss in an agent trace can denote either a local edit whose token state remains useful or a semantic rewrite that changes the answer; uniform near-match reuse conflates these cases. AdmitKV turns each miss into a block-level transfer decision: immutable provenance authorizes the source, monotone alignment and RoPE repair identify position-correct reusable KV, four deterministic micro-forward probes test computational agreement, and a type-conditioned empirical monitor can reject an otherwise eligible transfer. The production rule admits aligned code and tool state and sends document near-matches through exact lookup or recomputation. With the same exact-block-hash identity stage on SWE-bench Verified at 50% shared-file overlap, adding AdmitKV lowers median latency from 18.4 to 16.8 minutes (8.7%, 95% CI [5.3,11.8]%); relative to fresh execution it lowers latency by 22.9% (21.816.8 min, CI [19.6,26.3]%) and peak GPU memory by 30.6% (73.250.8 GB). Under a common validation-selected outcome envelope, its held-out point has 0.37% approximate false reuse and 0.7-point p99 event loss, versus 0.66–0.78% and 0.9–1.0 points for five selective-KV adaptations. Across the admitted-region perturbation grid, including the hardest 25% probe-aware setting, code/tool p99 loss is at most 0.9 point; the frozen thresholds also reduce p50 latency from 14.6 to 12.7 minutes on Natural Questions Open and from 8.7 to 7.3 minutes on GAIA while their paired quality intervals remain above the specified margins.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.