What Does Your Agent Remember? Auditing Storage, Reachability, and Behavioral Influence
Abstract
Information stored in an agent's memory may no longer support its later answers. We operationalize an audit that links durable records, active context, and behavior through native interventions and matched controls, returning explicit indeterminate verdicts when evidence or controls are unavailable. Across six agent-model configurations and 3,600 declared core executions, native Pi compaction reduces target adoption relative to restart by 20.8 points with DeepSeek V4 Flash and 15.0 points with GLM-5.3 Flash. Every success-to-failure pair in those contrasts still contains the complete source history in pre-probe durable records. Among the 25 GLM-5.3 Flash cases that succeed after restart and fail after compaction, seven, all recommendation tasks, still carry the literal target in active context. The procedure distinguishes preservation, inclusion, and use without treating any single surface as evidence of all three.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.