acceptodds
Under review as a conference paper at ICLR 2027

Experience Is Not Evidence: Coverage-Calibrated Memory Reuse for Tool-Using Agents

Abstract

Tool-using agents increasingly reuse their own past experience, yet a memory that looks relevant can quietly damage the trajectory it was meant to help. Having experience is not the same as having evidence that it transfers. We therefore replace the retrieval objective with an interventional one: the transfer gain of a request–memory pair, measured by running the same request with and without that memory, and generalized across pairs by tool-call structure rather than text similarity. What may be claimed about this gain is determined by coverage. Where practice contains no comparable structure, the sign of transfer is unidentifiable, making forced use-or-discard decisions provably lossy and motivating a genuinely resolving third action. We accordingly distinguish an average-value guarantee over candidate sets from a marginal per-pair certificate. The average-value guarantee opens on fresh confirmatory practice data and, with the policy frozen in advance, on a fresh fixed-design grid over held-out deployment domains. In contrast, independent validation with the representation frozen produces no pair-level openings and defers throughout. On a controlled benchmark designed so that the harmful candidate is the more textually similar one, only structurally incompatible memories are reliably harmful. Similarity and an LLM judge fail to identify them, whereas transfer-based gating incurs no observed harmful use on any tested model family and matches oracle utility on two of the three full-benchmark executors. Reliable harm appears only in our authored competing memories. A preregistered organic pilot produced isolated negative outcomes that failed its prespecified reliability gate and did not survive replication, bounding the scope of the result and reinforcing the need to measure transfer rather than infer it from a memory label.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.