acceptodds
Under review as a conference paper at ICLR 2027

What Did the Agent Learn From? Cross-Episode Memory Attribution for Self-Improving Agents

Abstract

Agentic attribution identifies which component of an LLM agent system caused an observed success or failure. This supports debugging, accountability, and targeted improvement. Existing methods are largely within-episode: they assign outcomes to agents, actions, tool calls, or prompt context in the current trace. Persistent memory agents break this assumption because the relevant cause may not appear in that trace. A memory written many episodes earlier can affect an outcome directly through retrieval or indirectly through memories written later. We introduce memory-level attribution for self-improving agents and formalise these two paths as retrieval influence and generation influence. If no allowed intervention retrieves a memory, retrieval-only data may contain no evidence that distinguishes it from an absent memory. In our deployed retrieval configuration, individually deleting the never-retrieved memories leaves every measured top-k list unchanged. Retrieval-only data therefore do not record effects passed through later memories in these cases. This limitation matters at scale. At the deployed retrieval budget, 87–92% of a 278-memory pool is never retrieved. We therefore propose a visibility-aware attribution framework that matches the attribution unit to what can be observed. MemAttr-V computes individual-memory Shapley values on the visible set, and MemAttr-B computes lineage-band Shapley values on the observed shadow. Across three environments and four agent LLMs, MemAttr-V’s correlation with measured leave-one-in utility is positive in every evaluated cell, while MemAttr-B’s mean correlation is positive in every environment within task splits. We release the memory pools, generation lineages, and measured targets as MemAttr-Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.