acceptodds
Under review as a conference paper at ICLR 2027

CAUSEWAY: Auditable Causal Memory for Long-Horizon LLM Agents

Abstract

Agent memory retrieves text that resembles the query, but an agent needs evidence that supports and influences the answer. Similarity and graph proximity can surface evidence that is stale, redundant, or confounded by a shared upstream cause, while retrieval alone does not establish whether an answer depends on that evidence. Destructive consolidation compounds this problem by deleting the history that temporal questions depend on. We present Causeway, a memory substrate that connects causal structure, evidence influence, and source auditing. Causeway compiles interaction history into a bitemporal causal graph whose typed edges become load-bearing after passing precedence, mechanism, and evidence gates. On its interventional path, each query is compiled into an explicit answer variable, and candidates are scored by contrasting their presence and absence in reader-visible evidence, with graph-selected background and a negative-control penalty. On the graph path, a source-linked proof supports answer generation and applicable removal/polarity checks; failed support checks trigger source replay or abstention. Non-destructive writes preserve earlier states for historical queries. Instantiated training-free with MiMo v2.5, Causeway achieves the highest Overall scores among evaluated methods on LoCoMo, LoCoMo-Plus, and MemoryAgentBench. Its strongest gains concern evidence localization, evolving states, and persistent constraints. Matched LoCoMo comparisons show that verification improves Answerable F1 by 6.30 points while reducing coverage by 1.29 points. On LoCoMo-Plus, targeted evidence use improves accuracy over Full Context with 76.99% fewer query-stage LLM tokens. Our code is available at [https://anonymous.4open.science/r/Causeway-3F66/](https://anonymous.4open.science/r/Causeway-3F66/)

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.