What Agent Memory Is Worth: A Zero-Cost Randomised Experiment Inside Test-Time Scaling
Abstract
Agent memory frameworks admit an item to the bank when an LLM judge calls the trajectory that produced it successful. Whether those items help anyone is unmeasured. Retrieval is not random, so a memory's presence is confounded with the difficulty of the tasks that retrieved it, and no observational score pulls the two apart. The confound can be removed for free. Memory-aware test-time scaling already runs N rollouts per task and hands each of them the same memories; we randomise which subset each rollout gets, which turns a budget already spent into an embedded randomised experiment. Each item's marginal causal effect then follows from a blocked difference of means over tasks, and memory governance becomes a decision rule on confidence bounds. We ran 2,450 rollouts across four WebArena sites and three conditions, 2,235 of which enter the paired comparisons. The per-memory effects come back small and bounded: 0.02 to 0.04 in magnitude, itself an upper bound, since the expected absolute estimate sits above zero under an exact null at these standard errors. Resolving effects that small would take roughly a thousand blocked observations per memory, between 28 and 225 times what a full campaign yields. Against the ReasoningBank + MaTTS baseline, our online arm spends 7,875 fewer input tokens per task (95% CI 355 to 15,395) with no detectable loss in ground-truth success; the data bound any such loss at 3.3 percentage points (+0.024, 95% CI -0.033 to +0.082). The frozen arm saves nothing and shows the same absence of a detectable difference (+0.017, -0.038 to +0.073). Neither result is an equivalence claim, and we declare no equivalence margin. The same design turns on the admission gate itself. ReasoningBank reports 72.7% judge accuracy against ground truth, AgentRewardBench 69.8% best-judge precision; neither reports discrimination. On a class-imbalanced rollout population, accuracy in that range is compatible with what we measure for a memory-stripped judge: AUC 0.538 (95% CI 0.514 to 0.562). Barely above chance. A composite outcome that weights such a judge at 40% flips the sign of its own arm comparison once the judge is dropped (-0.040 [-0.064, -0.016] against +0.040 [+0.010, +0.070], over 149 paired tasks), which is why every arm-level number above is reported on ground-truth success. We release the harness, the frozen task subset, and the per-rollout records.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.