MLPA: Multimodal Representation Learning via Masked Latent Prediction
Abstract
Integrating multi-modal Electronic Health Records (EHR) is difficult. Clinical streams vary widely in length, missingness is often clinically meaningful rather than random, and labeled outcomes are scarce. High-volume streams such as medication logs can dominate cross-modal attention and crowd out sparser signals, a failure mode we term modality washout. The objectives typically used for multi-modal integration add to the problem. Generative losses spend capacity on reconstructing high-entropy raw inputs, while supervised losses shape representations around a single downstream target. We introduce the Multi-modal Latent Predictive Architecture (MLPA), a self-supervised framework that learns patient representations from incomplete clinical records by predicting masked targets in uni-modal latent space. MLPA routes cross-modal exchange through a shared bottleneck transformer and represents missing modalities with learned empty-state tokens. We evaluate MLPA on 81,664 MIMIC-IV admissions and 188,094 eICU stays. On MIMIC-IV, MLPA matches a multi-modal masked autoencoder with 42.4% fewer parameters. On eICU, it achieves higher mortality AUPRC than the autoencoder (0.430 vs. 0.403). Without labels at pre-training, MLPA outperforms the semi-supervised method MUSE+ on in-hospital mortality in both datasets (eICU AUPRC 0.430 vs. 0.388). In eICU, this advantage holds for stays with one missing modality. MLPA also forms more compact clusters than MUSE+ for outcomes that neither model was trained on. An attention analysis shows that unconstrained cross-attention gives 47% of its attention to medication tokens, while the bottleneck keeps attention balanced across streams.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.