acceptodds
Under review as a conference paper at ICLR 2027

OT-MIMIC: Learning Medical Image Segmentation with Missing Modalities via Structure-Aware Optimal Transport

Abstract

Multimodal vision-language models (VLMs) for medical image segmentation often assume fully paired image-text data, an assumption rarely satisfied in real-world clinical datasets. In practice, large-scale medical corpora contain abundant partially observed samples and only a small subset of fully paired records. We study medical image segmentation under this incomplete multimodal setting and propose OT-MIMIC, a structure-aware transport-based completion framework that learns from partially observed data rather than discarding it. OT-MIMIC maps incomplete records to a dynamically refined memory bank and reconstructs missing representations or supervision through barycentric aggregation. We instantiate the framework with an asymmetric Fused Gromov-Wasserstein (FGW) objective in a CLIP-aligned latent space, combining semantic correspondence with relational consistency between observed-side and complementary-side structure. We additionally derive a high-probability excess-risk bound that decomposes learning error into empirical risk minimization, transport completion, recoverability, and function-approximation terms. Experiments on QaTa-COV19 and MosMedData+ show strong performance under both missing-text and missing-image-and-mask settings as complete-pair availability decreases from 50% to 1%, with the largest gains under extreme scarcity. Controlled ablations further show that the FGW instantiation improves over nearest-anchor retrieval and feature-mean completion under an otherwise fixed training pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.