When Memory Hurts: MemoryForge for Diagnosing Memory Utility in AI Agents
Abstract
Agent memory is usually evaluated by what it retains or retrieves, not by whether remembered experience improves the next task. We introduce MemoryForge, the first controlled evaluation framework, to our knowledge, that branches native agent memory from a shared frozen acquisition state to isolate the downstream behavioral effects of persistence, consolidation, full-history availability, and vetted content. These effects are instantiated as five task-paired interventions: Stateless, Frozen Context, Raw Memory, Dreamed Memory, and Oracle guidance. The protocol separately measures condition-independent artifact correctness and condition-aware process compliance, then records their joint attainment. On FinStateBench, three Claude tiers crossed with three native agent runtimes yield 1,350 executions over 30 held-out tasks. Oracle achieves the highest mean artifact score (0.971) but only 45.2% joint pass, compared with 72.6% for Stateless; Raw reaches 49.3%. Moreover, 58 of the 89 Stateless passes lost under Raw also fail under Oracle, exceeding the overlap expected under independence (44). Under Raw memory, runtime accounts for 78.0% of between-system variation, whereas model tier accounts for 3.1%. The five observed branches nevertheless cover 90.4% of cells, 17.8 points above the best fixed condition, with gains in all nine systems. These results expose a Memory Utility Gap: useful behaviors exist, but current agents do not reliably select, adapt, and execute the right memory mode for each task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.