acceptodds
Under review as a conference paper at ICLR 2027

Learning Consistent Multimodal Representations under Missing-Not-at-Random Shifts

Abstract

Missing modalities are ubiquitous in real-world multimodal applications. Although substantial progress has been made in learning from incomplete multimodal data, most existing methods implicitly assume that modality missingness occurs at random and is independent of the prediction target. In practice, however, modality availability can be correlated with the label, giving rise to missing-not-at-random (MNAR) phenomenon. Under such conditions, multimodal models may exploit missingness patterns as predictive shortcuts rather than learning the underlying multimodal semantics, leading to severe performance degradation when missingness patterns shift at inference time. We observe that such shortcut learning induces representation inconsistency, where the representation of the same instance changes with the set of available modalities. We theoretically show that the performance degradation under missingness-pattern shift is bounded by two factors: the discrepancy between source and target missingness patterns and the model's representation inconsistency. This analysis suggests that reducing representation inconsistency can improve robustness to missingness-pattern shifts without requiring prior knowledge of the target missingness distribution. Motivated by this insight, we propose AURA, a multimodal learning framework that explicitly encourages representation consistency across different modality-availability patterns. To enable comprehensive evaluation under MNAR conditions, we introduce Controlled MNAR Shift Evaluation protocol, that systematically controls both the dependence between missingness and labels and the degree of source-target missingness-pattern shift. Extensive experiments on multiple large-scale multimodal datasets demonstrate that AURA consistently improves robustness under MNAR shifts while maintaining strong performance on standard missing-modality benchmarks. On stroke type prediction, AURA's relative improvement in balanced accuracy over the strongest competing method increases from 20.4% under milder MNAR shifts to 60.2% under more severe shifts. Our code is available at https://anonymous.4open.science/r/Learning-Consistent-Multimodal-Representations-under-Missing-Not-at-Random-Shifts

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.