M³DT-RL: Multimodal Multidisciplinary Diagnostic Teaming with Reinforcement Learning
Abstract
Multidisciplinary consultation integrates complementary specialty perspectives, but training a shared policy for its distinct diagnostic roles remains challenging. We study reinforcement learning over the full multidisciplinary diagnostic trajectory, termed *MDT-RL*, for multimodal differential diagnosis (DDx) reranking. MDT-RL faces two challenges: trajectory-level rewards provide coarse credit across stages, while uneven learning progress makes fixed rollout allocation inefficient. We introduce **M³DT-RL**, combining *Modular Stage Training* (MST) and *Mastery-Driven Sampling* (MDS). MST jointly optimizes full MDT trajectories and isolated Initial, Expert, and Final tasks with revision-aware credit, while MDS tracks task mastery and adaptively reallocates rollouts toward insufficiently learned tasks. We further construct **M²DDx**, containing 4,225 acute and emergency-related cases and 22,208 images, with 537 cases published from March 2026 onward reserved for temporal holdout evaluation. Across six multimodal DDx benchmarks, **M³DT-RL-4B** achieves macro-average Hit@1/MRR of **60.89%/75.62%**, and **M³DT-RL-9B** reaches **62.93%/77.12%**. Both outperform the compared commercial MLLMs, larger open general and medical MLLMs, and prompted multi-agent workflows. These results show that multidisciplinary diagnosis depends not only on model scale or agent count, but also on how diagnostic roles and their coordination are trained. Our code is available at: https://anonymous.4open.science/r/M3DT-RL-2460/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.