acceptodds
Under review as a conference paper at ICLR 2027

MedOmni: Benchmarking Omni Modality Integration in Medical Agentic Systems

Abstract

Clinical decisions require reasoning over heterogeneous medical evidence, yet existing benchmarks typically restrict each task to a small set of predefined modalities. This leaves medical agents' ability to select, access, and integrate broad evidence insufficiently evaluated. We introduce MedOmni, a clinically grounded benchmark developed with a multidisciplinary clinical consortium. To our knowledge, under our counting criterion, MedOmni provides the broadest modality coverage per task and is the first benchmark designed specifically for omni-modal medical agent evaluation. Its specialised infrastructure decouples tasks, modalities, and encoders through an extensible modality perception service, allowing new modalities to be incorporated without redesigning existing tasks. A structured assurance mechanism connects benchmark requirements to auditable checks across data, perception, tasks, and evaluation runs. Experiments reveal substantial differences in how agents benefit from and reason over heterogeneous evidence, identifying evidence selection and integration as central challenges for omni-modal medical intelligence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.