acceptodds
Under review as a conference paper at ICLR 2027

MAGMA: Multimodal Agentic Medical AI with Context-Based Iterative Refinement

Abstract

Predictive machine learning models hold substantial promise in healthcare, yet their development typically requires significant domain expertise and careful methodological reporting. Recent agentic AI systems have demonstrated strong coding capabilities, including for end-to-end machine learning pipeline construction. However, their use in healthcare raises additional requirements since workflows must be clinically-grounded, transparent, and reported in sufficient detail to support appraisal and reproducibility. We introduce MAGMA (Multimodal AGentic Medical AI), an agentic system for end-to-end clinical model development and evaluation. MAGMA's central contribution is a judge-driven iterative refinement loop in which a judge agent audits each attempt against a rubric derived from established clinical prediction modelling guidelines and returns a composite refinement score combining predictive performance with reporting compliance. This score and the item-level feedback guide the execution of the next attempt, so that clinical reporting compliance drives refinement. We also constructed MAGMA-Suite, a new benchmark of 10 clinical case studies spanning structured electronic health records, CXR images, ECG signals, and clinical text. We also propose MAGMAScore, a guideline-derived 22-item metric for assessing the clinical grounding of autonomous modeling workflows. All systems are evaluated on a fixed, patient-disjoint held-out partition across four prompt templates and three independent runs per configuration, and results are reported on a completion-weighted basis that accounts for failed runs. Under this protocol, single-agent baselines complete only 22.5% of attempted runs, while MAGMA completes every attempted run, and attains the highest completion-weighted AUROC in 9 of the 10 case studies. In a blinded expert adjudication conducted by two clinical machine-learning assessors against a rubric held out from the refinement loop, MAGMA received the highest score in 8 of 10 case studies (mean 0.74 versus 0.62 across baseline packages). We release the system, benchmark, and evaluation pipelines to support reproducible research on autonomous machine learning in healthcare.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.