Amortizing Causal Memory Verification via Evidence Reuse in Long-Horizon LLM Agents
Abstract
Long-horizon large language model (LLM)-based agents increasingly rely on external memory, but relevance-based retrieval does not establish whether memory exposure improves agent behavior. Causal Memory Intervention (CMI) provides a stronger behavioral criterion through matched interventions, but applying Full-CMI to every candidate incurs repeated verification costs. We study whether this cost can be amortized across sequential memory decisions. We propose the Causal Evaluation Reuse Policy (CERP), a selective verification policy that fits a rejector with history-conditioned features on an initial Full-CMI-evaluated prefix and applies the frozen policy to subsequent candidates, rejecting candidates before utility verification. On Causal-LoCoMo, CERP achieves 12.7% net call savings relative to Always-Verify after accounting for the initial Full-CMI cost, with a pooled rejection error rate of 8.8%, a 0.02 decrease in the extractive task proxy, and a 0.051 decrease in Useful-Memory F1. Matched random rejection and direct nearest-label reuse incur higher rejection error (16.3-23.7% vs.8.8%). Budget and horizon analyses characterize when the initial verification cost is recovered.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.