acceptodds
Under review as a conference paper at ICLR 2027

Amortizing Causal Memory Verification via Evidence Reuse in Long-Horizon LLM Agents

Abstract

Long-horizon large language model (LLM)-based agents increasingly rely on external memory, but relevance-based retrieval does not establish whether memory exposure improves agent behavior. Causal Memory Intervention (CMI) provides a stronger behavioral criterion through matched interventions, but applying Full-CMI to every candidate incurs repeated verification costs. We study whether this cost can be amortized across sequential memory decisions. We propose the Causal Evaluation Reuse Policy (CERP), a selective verification policy that fits a rejector with history-conditioned features on an initial Full-CMI-evaluated prefix and applies the frozen policy to subsequent candidates, rejecting candidates before utility verification. On Causal-LoCoMo, CERP achieves 12.7% net call savings relative to Always-Verify after accounting for the initial Full-CMI cost, with a pooled rejection error rate of 8.8%, a 0.02 decrease in the extractive task proxy, and a 0.051 decrease in Useful-Memory F1. Matched random rejection and direct nearest-label reuse incur higher rejection error (16.3-23.7% vs.8.8%). Budget and horizon analyses characterize when the initial verification cost is recovered.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.