DriftPair: Counterfactual Carry Evaluation Reveals Memory-System Reversals Under Cross-Session Drift
Abstract
Independent-session scores can favor memory architectures that deteriorate when stale state crosses session boundaries. DriftPair exposes this selection error with a branch-aligned protocol that crosses access to pre-session memory with a shared revision realization while matching tasks, policies, models, tools, decoding, budgets, and exogenous randomness. Across 990 tasks in four verifier-backed tool-agent families, carry benefits reuse-oriented policies in stable environments, yet revision produces 4.7- and 5.2-point carry–drift interactions for tiered and rolling summaries, compared with 1.3 and 1.2 points for temporal and verify–evict memory. These interaction gaps change architecture choice: independent sessions rank MemGPT/Letta first, whereas DriftPair ranks it fourth; carrying that choice into the revision regime costs 2.4 session-20 success points, adds 3.3 interference points, and consumes 2,070 additional tokens per episode. A Qwen2.5-72B-Instruct replication reproduces the demotion. By making pre-session state the controlled variable, DriftPair links memory mechanisms to an architecture-selection consequence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.