DiagMem: Closing the Diagnostic Memory Gap for Evidence-Grounded Multi-Agent Rare Disease Reasoning
Abstract
Rare disease diagnosis often depends on evidence scattered across symptoms, laboratory results, family history, phenotypes and external disease knowledge. Current LLM-based diagnostic systems usually treat this process as prompt-level prediction. They may produce plausible disease rankings, but they do not maintain what has been established, which hypotheses have been considered, what remains uncertain or how new evidence changes the case. We introduce Patient Diagnostic Memory, a shared-memory framework for multi-agent rare disease diagnosis. For each patient, the framework maintains a diagnostic state, event records and semantic memory entries that store reusable clinical evidence. Specialized agents extract clinical facts, normalize phenotypes with Human Phenotype Ontology terms, interpret laboratory findings, analyze family history and retrieve external disease knowledge. A diagnostic controller reads the evolving memory to decide which evidence source to use next and when to produce a ranked diagnosis. We evaluate the framework on MIMIC-IV-Note and RareArena with four backbone LLMs. Memory-augmented diagnosis improves top-ranked accuracy across all models, with larger gains on rare disease cases. Average top-1 accuracy increases by 14.5 percentage points on RareArena and 4.3 points on MIMIC-IV-Note. These results suggest that patient-level diagnostic memory can make multi-agent clinical reasoning more traceable, modular and evidence-grounded.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.