acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Failure Regimes in LLM Agent Memory with Extraction-Pass Disagreement

Abstract

LLM agents build long-term memory by consolidating interaction history into a structured store, and this consolidation is lossy in two ways. In the first, a stochastic pass drops information that another pass may keep, and computation removes the error. In the second, a capability-limited backbone compresses history into a low-rank sketch, and the bias it introduces is beyond the reach of computation. The two failures look the same in aggregate accuracy, so existing systems apply the same remedy to both. We propose Regime-Routed Construction (RRC), a training-free configuration layer whose routing follows the diagnosed failure regime. RRC repeats consolidation on the same history and reads how much the passes disagree. The disagreement estimator, combined with backbone capability, routes the vote budget, the retrieval depth, and the answer policy. A noisy reading brings more extraction passes merged by majority vote. A quiet reading deepens retrieval and calibrates the answer policy when the backbone sits above its capability threshold, and declines what the store cannot support below it. A Fano-type argument shows that below a capability threshold, error on relational questions is irreducible. Across two benchmarks and backbones from 1.5B to frontier API models, RRC improves on the strongest published LoCoMo result by +4.7 F1 and outperforms the RL-trained Memory-R1 at matched backbones without gradient updates. On LongMemEval, the gains are small and insignificant, consistent with a diagnosis of little headroom.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.