Understanding and Mitigating Missing-Modality Effects in Medical MLLMs
Abstract
Missing modalities are common in medical settings and may hinder clinical decision-making. While prior work has explored ways to mitigate this issue, its impact on medical multimodal large language models (MLLMs) remains underexplored. This raises two questions: (1) how existing medical MLLMs perform under missing modalities; (2) how to improve their performance when modalities are missing. To address the first, we introduce MedGap, a MIMIC-IV-based benchmark with CXR, ECG, and laboratory data to systematically evaluate MLLMs under missing modalities. Unlike prior benchmarks that primarily focus on arbitrary modality combinations, MedGap further examines missingness within disease-relevant pairs, avoiding interference from irrelevant modalities. Our evaluation shows that more modalities do not automatically translate into better performance, while the paired diagnosis evaluation further shows that the impact of a missing modality depends on its relevance to the diagnosis target. Motivated by these findings, we address the second question with MIRI, a two-stage framework that combines probabilistic cross-modal recovery in latent space with teacher-guided integration of the recovered information into the LLM. Experiments on MedGap show that MIRI improves average performance across missing-modality settings and multiple model backbones. Further analysis shows that cross-modal recovery captures missing-modality information and downstream integration enables its effective use, with recovery becoming more beneficial as fewer modalities remain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.