Separating Laboratory Values and Care Events in ICU Representation Learning
Abstract
Self-supervised models of health records typically merge laboratory results and care events (laboratory orders and medication starts) into a single sequence read by one transformer, mixing signals about the patient’s physiology with signals about the care process. Because care processes differ across hospitals, we study whether these two types of information are better represented separately. We propose a coupled model in which laboratory values and care events are encoded by separate encoders and aligned with a contrastive loss. Across three ICU databases (eICU, MIMIC-IV, HiRID), ten seeds, and eight downstream tasks, the coupled model achieves significantly higher AUROC than the single-stream baseline in 67 of 85 comparisons and significantly lower AUROC in only one, with a mean improvement of +0.019, including at held-out sites and after cross-database transfer. The coupled model also preserves more information about measured laboratory values: linear probes recover laboratory values substantially better from its representation, with R² gains of +0.042 to +0.108 in distribution, +0.053 and +0.072 at held-out sites, and +0.075 to +0.238 after transfer. Ablations show that separating the encoders, using a set encoder for laboratory values, and adding a laboratory–vital contrastive loss improve laboratory-value recovery, whereas the cross-encoder contrastive loss primarily improves downstream prediction and can reduce recovery on one database. These results suggest that separately encoding laboratory values and care events, while aligning their representations contrastively, can improve both predictive performance and transfer of ICU representations across sites and databases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.