CARE-Fuse: Reliability-Aware Multimodal Fusion Under Missing and Conflicting Clinical Evidence
Abstract
Multimodal fusion must remain effective when evidence is missing or expert predictions conflict. We introduce CARE-Fuse, a task-adaptive late-fusion framework that incorporates evidence availability and cross-expert disagreement into validation-trained gates. A reliability gate is trained with label-independent validation perturbations to improve robustness to missing and conflicting inputs. We evaluate CARE-Fuse on clinical prediction tasks by combining first-24-hour vital signs and laboratory features, temporal representations, report embeddings, and language-model scores. On a locked test cohort of 14,637 MIMIC ICU stays, CARE-Fuse achieves AUROC/AUPRC of 0.8674/0.5344 for in-hospital mortality and 0.8310/0.6878 for stays longer than three days. Relative to the strongest single expert, paired bootstrap analysis with 5,000 draws yields gains of 0.0428 AUROC and 0.0671 AUPRC for mortality, and 0.0386 AUROC and 0.0960 AUPRC for prolonged stay; all 95% confidence intervals exclude zero. Under missing temporal or report evidence and conflicting language-model outputs, the perturbation-trained reliability gate reduces degradation compared with an unaugmented gate, while a clinical-only stack remains competitive on clean data. Calibration and selective-prediction analyses further characterize evidence coverage, expert disagreement, and case-review behavior. These findings support perturbation-trained fusion with explicit availability and disagreement features as a practical approach to robustness in multimodal clinical prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.