acceptodds
Under review as a conference paper at ICLR 2027

MedMME: Benchmarking Long-Horizon Medical Memory for Multimodal Agents

Abstract

Long-term memory is a critical capability for multimodal medical agents to integrate evolving clinical information throughout extended interactions. However, existing medical memory benchmarks cover a limited range of modalities and tasks, failing to fully capture the multimodal and dynamic nature of clinical information. To bridge this gap, we introduce MedMME, a benchmark for evaluating long-horizon medical memory in multimodal agents. MedMME integrates datasets across 10 clinical domains and 9 medical visual modalities, covering two medical settings: multimodal clinical dialogue and medical procedure videos. We divide clinical trajectories into sessions to evaluate agent memory under a streaming input protocol. The benchmark also introduces physician-validated clinical memory to evaluate memory extraction, updating, and retrieval. We systematically evaluate 19 memory-augmented multimodal agents built on nine memory mechanisms. We find that: (i) agent performance declines as multimodal clinical interactions progress; (ii) existing memory mechanisms perform unevenly across medical visual modalities; and (iii) agents with medically adapted backbones perform better, while backbone models scaling alone yields limited gains. By making these memory capability gaps measurable, MedMME supports the evaluation of future multimodal agents for long-horizon medical interactions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.