LifeMemBench: Benchmarking Long-Term Memory for Evolving User States through Personalized Planning
Abstract
Personalized language-model agents increasingly interact with users over extended periods, yet user preferences, constraints, goals, and circumstances evolve over time. Agents must therefore recover relevant information, resolve what remains valid, and use the resulting user state in subsequent decisions. Existing evalua- tions have studied long-term retrieval, temporal reasoning, memory updates, and preference following, but often separately from the downstream decisions they are intended to support. We introduce LifeMemBench, a benchmark for longitudinal personalization under evolving user states. LifeMemBench constructs controlled latent user-state trajectories before realizing them as chronological interactions, providing query-time supervision for both current and historical states. It jointly evaluates state recovery, personalized planning, and replanning through a common process of recovering relevant information, resolving the valid state, and acting on it. Spanning 8K–64K tokens, planning and replanning are evaluated against explicit user-specific and feasibility constraints to test whether longitudinal information af- fects downstream behavior. We further introduce CoMem, a lightweight state-aware memory system that preserves historical evidence, tracks state changes, and derives an actionable current-state view. Under a shared Qwen3.5-4B backbone, CoMem achieves 47.23% QA correctness and 46.53%/41.24% on planning/replanning, compared with the strongest competing results of 40.27% and 33.05%/29.55%. Across systems and controlled ablations, we find that stronger information recovery does not necessarily yield better personalization: successful long-term interaction requires not only remembering user information, but resolving the right state and consistently translating it into action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.