Evaluating and Understanding Agent Memory Systems for Chronic Care
Abstract
Memory is essential for large language model (LLM) agents in chronic care, where patient histories span years of visits across multiple departments. Many memory systems have been proposed for LLM agents, yet existing benchmarks, drawn from general-domain dialogues or medical records over limited years of visits and departments, cannot reveal how they behave or why they fail in chronic care. To fill this gap, we introduce MEME-Bench, a clinically grounded benchmark of patient histories synthesized from real chronic-care records, with records spanning over a decade, patients up to 1,718 visits, and 1,646 questions on eight capabilities asked as the history accumulates. Leveraging MEME-Bench, we evaluate 11 memory systems and a long-context baseline across seven backbones and uncover three critical findings: (1) no memory system dominates across capabilities, and none outperforms the long-context baseline on every axis; (2) longitudinal trajectory tracking and cross-department integration remain poor for all systems; and (3) tracing the required facts through the memory pipeline reveals retrieval as the most common source of failure, rooted in evidence scattered across departments, recurring measurements and precise medical terminology. Controlled ablations further show that combining agentic retrieval with date-filtered search improves both of the hardest axes, whereas prompting cannot replace these components.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.