PERSIST-Med: Evaluating Implicit Use of Patient History in LLM Memory Systems
Abstract
Patient-facing medical assistants must use earlier patient facts even when a new question does not mention them. A past drug reaction, for example, may still rule out a treatment today. Memory benchmarks largely test factual recall and updates, but rarely ask whether stored history changes a clinical decision unprompted. We introduce PERSIST-Med, a benchmark of 1,422 questions over 200 synthetic patient histories spanning 540 days. It distinguishes constraints that remain binding from values superseded by later evidence, and tests implicit application, update discrimination, conflict resolution, and abstention. Reference answers are specified from a guideline-grounded fact base before dialogues are generated. On a 40-patient evaluation set, two frontier models apply critical facts in 93.1–96.9% of implicit-application items with full history, versus 73.1–76.2% with mem0 or A-MEM, with about five times as many contraindicated recommendations. The systems fail at different stages: mem0 often fails to retrieve a fact it stored, while A-MEM often fails on items whose retrieved notes already contain the fact's text. On the full 200-patient corpus, adding a short, query-blind patient summary to mem0 reaches 94.1% on 797 implicit-application items, compared with 95.1% for full history. On 296 update items, however, its stale-value error rate remains 29.4%, compared with 1.0% for full history.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.