From Bottleneck Labels to Repair Surfaces: Factorial Diagnosis of Agent Memory
Abstract
An oracle repair may recover many tasks from a native memory-agent pipeline, yet tell us little about where the original failure occurred if another repair recovers the same tasks. We examine this ambiguity by crossing repairs to upstream state, retrieval, and read context. In the primary setting, upstream-state and read-context repairs each add about 91 pp from the native cell, but no more than +0.45 pp once the other is active; they correct 199 of the same 206 native errors. Across all seven complete surfaces, the two repairs have negative accuracy interactions. Retrieval repair has smaller native gains across eleven settings, and paired intervals do not resolve a change in its gain after upstream repair. The same trajectories yield nominal p < .05 retrieval-background contrasts with individual judges, but none with adjudicated labels. Stage-level interpretation therefore depends on both the repair background and the scoring rule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.