Transformer-Based Multimodal Retrieval with Knowledge Graph Supervision for Clinical Decision Support
Abstract
Diagnostic decision support that combines imaging, unstructured clinical narra- tives, and structured tabular measurements is increasingly bottlenecked by threeproblems: (1) representations that ignore relational biomedical knowledge; (2) re-trieval mechanisms that are dominated by lexical or dense but semantically shallow matches; and (3) evidence sets that are redundant and unable to explain individual predictions. This paper introduces a knowledge-graph-supervised multimodal re-trieval transformer for diagnostic decision support whose primary methodological contribution is a coverage approximation guarantee. It jointly aligns imaging, textual, and tabular encoders under a unified biomedical ontology through a graph-aware supervision loss. The system retrieves candidate patient modal coverage, andfactual triples via a cross-modal attention transformer, and compresses the resulting candidate pool through a submodular utility, which jointly rewards diagnosticrelevance, cross-modal coverage, and ontological consistency while penalizing redundancy. Extensive experiments on MIMIC-IV coupled with MIMIC-CXR,TCGA multi-omic imaging cohorts, and ADNI multimodal Alzheimer data, usingthe Unified Medical Language System as the reference biomedical knowledge graph, show improvement over the strongest dataset-specific baseline of approximately 93.4% macro AUC and 87.6% macro F1, reducing average evidence-set size by 41.7% (24.7 to 14.4 facts), inference latency by 45.3% (342 to 187 ms),and expected calibration error by 47.5% (6.1% to 3.2%). Ablation studies and robustness tests under 30% modality dropout confirm that both the knowledge graph supervision loss and the submodular selection stage contribute independently and complementarily to these gains, positioning this model as a strong foundation for radiology triage and ward-level clinical decision support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.