acceptodds
Under review as a conference paper at ICLR 2027

MedMDT: Self-Evolving Multidisciplinary Agent Teams for Clinical Reasoning

Abstract

Clinical reasoning over electronic health records (EHRs) requires integrating heterogeneous longitudinal evidence, resolving competing hypotheses, and learning from experience without observing future events. We present MedMDT (medical multidisciplinary agent teams), an event-driven system that couples task-routed specialists with a coordinating chief and an evidence critic. Its central design is a dual-state adaptation protocol: Decentralized Low-Rank Policy Optimization (D-LRPO) learns role-specific parametric policies from reward-scored rollouts, whereas Critic-Bandit Skill Policy Optimization (CB-SPO) selects, validates, and versions reusable textual policies. To eliminate data leakage, the architecture isolates adaptation, inference, and evaluation through strict temporal filtering, hidden labels, and frozen benchmarks. We evaluate MedMDT on MedEHR, comprising 5100 task instances spanning common, rare, cross-dataset, and randomized clinical trial (RCT) settings. With a shared chief backbone, MedMDT attains 0.4730 Sample-F1 on MedEHR-Common, compared with 0.4108 for the strongest same-chief baseline. MedMDT Pro achieves the highest MedEHR-RCT average among the compared systems (0.6530) while remaining competitive on the other subsets. Increasing adaptation data from 600 to 3000 cases raises Sample-F1 from 0.4730 to 0.5699 on MedEHR-Common and from 0.6126 to 0.6532 on MedEHR-RCT. Component ablations and prediction-transition analyses support complementary contributions from specialization, critique, and the two adaptation channels. Code will be released on GitHub.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.