Experience Memory and Episode Completion: A Pre-Registered Audit on Two LLM Agents
Abstract
Experience memory, a bank of distilled past episodes retrieved into an agent's context, is widely reported to improve LLM agents, but outcome-bearing memories may also induce premature commitment. We pre-registered a 33-arm dissection of an ExpeL-style memory bank on three multi-hop QA benchmarks (about 29,700 logged episodes, plus 3,000 in two later blocks): the design was hash-frozen before any test episode, four rounds of numeric go/no-go gates ran without a threshold change, and every deviation is reported (Appendix A). For Qwen3.5-27B, the pre-registered NoMem–Mem contrast is equivalent within the registered margin (paired ΔEM +0.6pt; TOST 90% CI [−1.2, +2.4] within ±4pt). None of six content manipulations produces a significant change after Holm correction, and memory does not shorten search: unconditional searches rise, and episodes that reach an answer (a set the treatment itself shifts) use the same number with and without the bank. Memory instead changes how episodes end: the share that exhausts the call budget without answering rises from 14.0% to 25.2%. A post-hoc, single-seed ablation shows that episode-completion mechanics alone (a budget extension and a terminal forced finish) reproduce most of the gain of the evidence-citation instruction arms in three of four decompositions: mechanics recover +6.2pt and the instruction adds +2.0pt beyond them (90% CI [−0.2, +4.2]). Two blocks on a second hardware setup (4×A6000) give, for mechanics and then instruction, +3.8 and +0.4pt with the bank and +2.2 and +2.6pt without it. At matched mechanics the memory contrast is +2.4pt on both hardware setups, too imprecise to establish either a nonzero effect or equivalence. Hard enforcement does not outperform the instruction, a content-free bookkeeping instruction reaches a similar EM point estimate, and optimizer-evolved instructions (GEPA, MIPROv2) do not significantly outperform the hand-written instruction on the full test set; their minibatch gains do not transfer. Two conclusions reverse on gpt-oss-120b, a mixture-of-experts agent with lower base EM on the matched subset: memory helps (+8.5pt, p = 3.7 × 10⁻⁷, two seeds) and stored answers carry the value (redaction −5.3pt). Memory-bank effects differed across the two agents studied here and should be reported per agent. For Qwen3.5-27B, completion mechanics plus one instruction, without a bank, trails no memory configuration by more than 3pt (pooled point estimates), at under two-thirds of the tokens per episode of Mem+Cite.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.