acceptodds
Under review as a conference paper at ICLR 2027

DELETECERT: Counterfactual Certificates for Deletion in Persistent Agent Memory

Abstract

Deleting a record from an agent's memory is not the same as deleting its influence. A persistent language agent turns one user artifact into summaries, embeddings, graph edges, profiles, skills, cached plans, messages to other agents and sometimes parameter updates; removing the original row leaves this derived state in place. We define deletion counterfactually: an artifact is deleted, on a declared probe law, when the deployed agent behaves like the agent of the world that never observed , and residual influence is the excess risk over that never-seen agent. Around this definition we build \method: a typed artifact-lineage graph, a closure that deletes, recomputes or keeps each descendant under a risk budget and a utility floor, and a scoped audit that bounds excess risk under the declared law and reports whether the bound meets the target. We prove that faithful recomputation reproduces the never-seen world exactly and degrades gracefully in total variation, that the audit's bound is valid for a fixed post-deletion state, that minimum-cost closure is NP-hard, and that no finite probe set certifies worst-case equivalence. On a synthetic structural benchmark, deleting the raw record leaves excess risk; the audit's bound held on all audited cases, of which report a failed deletion, and the greedy planner found a plan on of cases, while the exact optimum of its own problem exists on all of them. On memory written by a language model, the closure removes a record held in one session but loses unrelated content it deletes rather than recomputes, and misses a record stated across two sessions until its detector reads the parts and its scope covers every root; the audit reports each failure. The benchmarks are a structural simulator and a small language-model world, and every number describes them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.