Same Curve, Different Regimes: A Longitudinal Audit of Memory-Augmented Agents
Abstract
Memory-augmented agents adapt after deployment by accumulating interaction experience, and their aggregate success often rises before flattening. Does a plateau indicate a stable capability gain? Similar curves can conceal a stable solved-task set, gains offset by losses, changes at the scale of repeated execution, or a task bank with little remaining discriminative range. We propose a longitudinal audit of frozen memory states that jointly examines across-seed coverage, matched reruns of the same state, and the item-level composition of change. On a fixed ALFWorld bank, three backbone operating points and two memory mechanisms reveal gain–loss cancellation, repeat-scale change, and ceiling-limited evidence. On a smaller ScienceWorld panel, two further backbones show different early trajectories and later changes near the repeat scale. A plateau should therefore prompt this audit before it supports a claim of convergence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.