acceptodds
Under review as a conference paper at ICLR 2027

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

Abstract

Modeling user memory is essential for building personalized AI agents. However, existing benchmarks mainly focus on sparse user information explicitly stated in conversational narratives. In contrast, real-world information is scattered across diverse digital sources, such as messages, calendars, and photos. Such observations form a dense and evolving data stream that only indirectly reveals underlying user preferences and habits. This mismatch limits the effective evaluation of agent memory systems. However, collecting real user data remains prohibited due to privacy concerns. This paper introduces LifeBench, a simulation framework that combines *hierarchical outline planning* for consistency control with *Objective Agents* for grounding generated data in real-world information. Evaluation shows that LifeBench better captures real-world preference shifts than existing memory benchmarks. In addition, its Objective-Agent mechanism enables digital traces that resemble human mobility patterns in real-world data without supervised training. We further synthesize a moderate-sized dataset and evaluate state-of-the-art memory systems. Results show that LifeBench more clearly exposes performance gaps among current systems. Deeper analysis reveals insights across data distributions and settings, including, e.g., *sparse vs dense data distributions*.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.