What to Learn and What to Trust: Multimodal Learning with Reliable Supervision for Bioactivity Prediction
Abstract
Cell Painting phenotypes and molecular structures provide complementary views of compound activity, but learning from them is complicated by replicate variation, non-equivalent structure–phenotype relationships, and sparse assay supervision. Existing multimodal approaches typically align paired molecular and morphological observations, implicitly encouraging agreement between the two modality-specific representations. We propose BioCoReT, a two-stage framework that addresses what to learn from rich multimodal observations and what to trust under sparse assay supervision. First, BioCoReT performs contrastive learning across fused views of the same compound. It constructs biology-aware multimodal views from matched controls and Cell Painting replicates, and learns consistency across these compound-level views while preserving modality-specific information. Building on these representations, assay-specific specialists learn task-specific decision boundaries from the available labels and generate candidate predictions for unlabeled compound-task pairs. BioCoReT then evaluates each candidate using complementary evidence from teacher confidence, molecular and morphological topology, and replicate consistency, transferring only reliable task-specific knowledge to a unified multi-task student. BioCoReT consistently outperforms existing methods across three public Cell Painting benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.