Towards Robust Long-Lived Agents: Benchmarking Experience Lifecycles in Stateful Environments
Abstract
Long-lived agents must discover how an unfamiliar environment works and keep the resulting experience useful as conditions change. Existing benchmarks have expanded from historical information retrieval to interactive experience reuse and adaptation, yet maintaining experience under costly exploration and unannounced changes remains underexplored. We introduce ReScopeBench, a benchmark connecting experience acquisition, reuse, invalidation, and revision across 18 interactive environments. Our experience-first construction organizes 16-task streams around structural experience about resource relationships and rule experience about local operating conditions. Agents pursue varied goals under explicit exploration budgets, while silent changes invalidate selected dependencies and preserve others. Paired evaluations with and without retained experience measure benefits, disruption, and recovery through objective outcome scoring. Experiments with seven models show that experience acquired through interaction improves subsequent task performance, but silent environmental changes erode these gains and recovery varies substantially across models. Trace analyses further reveal a gap between retaining valid experience and applying it effectively. ReScopeBench makes maintaining actionable experience an explicit target for long-lived agent evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.