LIFEFORGE: SYNTHETIC DIGITAL LIVES FOR MULTI-SOURCE AGENT MEMORY EVALUATION
Abstract
AI assistants increasingly build memory from a user's personal data spread across many sources rather than from conversation alone. Evaluating them requires multi-source personal data with verifiable ground truth, which real users cannot provide for privacy reasons, and existing memory benchmarks remain almost exclusively conversational. We present LifeForge, a two-stage framework that synthesizes a virtual user's multi-source data and compiles memory questions from it without human authoring. Stage one derives every record from one hidden event timeline, projecting each event into mutually consistent calendar entries, emails, office files and photos. Because every record traces back to a known event, Stage two can select the evidence for a question and keep only questions whose answer is unique. We release 11 digital lives in six countries and four languages, 29K records in all, and 1,034 questions covering four memory operations. Scoring five open-source memory systems under one harness, on answer quality, record-level retrieval and the cost of writing and reading memory, exposes two shared failures. They retrieve records one at a time, so most collapse once an answer has to gather many records from one stretch of time; giving the same reader the gold evidence closes most of that gap, which puts the fault in retrieval. They extract facts at ingestion and discard the details questions ask about, a loss a stronger reader widens rather than repairs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.