TRACE: Recall as Execution for Accountable Memory
Abstract
Agents that work across sessions must remember what matters, apply updates, and honor deletion and access requests; retrieved text cannot say whether a fact is still valid or permitted. We introduce TRACE, which implements Recall-as-Execution: each write records an operation (add, replace, delete), and each recall executes them to decide which facts hold at the requested time for the requesting principal. Every structured answer carries a certificate: the instructions it rests on and why each withheld fact was excluded. Temporal queries, forgetting and counterfactual replay are one call with a different time or ledger. We prove execution deterministic and grouping-independent, and check it against the deployed executor. On fixed ledgers all 659 recalls are admissible with no misuse, against 0.238–0.826 for a dense reader, and all 435 evaluated certificates verify, against 0.000 for baselines reporting no exclusions. On AutoMemoryBench, TRACE satisfies task success, safety and temporal validity together on 4.0% of cases, against at most 0.1% for any full-lifecycle vendor. A matched control locates the margin: folding the same operations into present state matches our admissibility, so the gain is in executing updates, not our format; but present state alone caps time-indexed queries at 0.333, where executing the ledger answers 1.000. Ablating supersession or governance collapses safety to 0.253 and 0.000 while task accuracy holds near 0.92: the result is in the execution semantics, not the backbone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.