KA-MIL: Knowledge-Anchored Multiple Instance Learning for Patient Classification from Single-Cell Transcriptomics
Abstract
Patient classification from single-cell transcriptomics requires aggregating highly heterogeneous cell populations under limited patient-level supervision. Standard data-driven multiple instance learning (MIL) pooling lacks an explicit mechanism for organizing cellular contributions by biological concepts. We introduce Knowledge-Anchored Multiple Instance Learning (KA-MIL), an architecture that utilizes embeddings of biological descriptions as semantic queries for attention-based cell aggregation. Within this framework, a geometry regularizer aligns the pairwise similarities of projected semantic queries with those of the original text embeddings, while learnable residual queries provide complementary task-adaptive capacity. Using Gene Ontology (GO) descriptions with cohort-wide, label-free vocabulary construction, KA-MIL achieves the highest mean macro-F1 in all five primary evaluation settings: four clinical cohorts and a reduced-training setting. On the COVID-19 cohort, it improves macro-F1 over Hier-MIL by 7.04 percentage points. Beyond prediction, our COVID-19 analysis reveals a recurring relative humoral–myeloid attention contrast between plasmablast and CD14 monocyte summaries in 54 of 58 higher-severity test patients across five folds. Together, these findings demonstrate that knowledge-anchored aggregation effectively couples accurate patient classification with the structured, description-guided inspection of cellular summaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.