Multimodal Learning with Missing Modalities via Multi-Level Alignment for Drug–Target Affinity Prediction
Abstract
Drug–target affinity (DTA) prediction estimates the binding strength between a small-molecule drug and a target protein. Existing methods typically rely on either sequence or structure information, limiting the joint exploitation of complementary supervision from heterogeneous datasets. Label heterogeneity is often overlooked, leading to the conflation of distinct experimental measurements into a single prediction target. To address these issues, we propose M3A-DTA, a missing-modality, multimodal, and multi-assay framework that learns from any available subset of drug sequence, protein sequence, ligand structure, and protein structure. Modality-specific pretrained encoders provide chemically and biologically informed representations. Through multimodal alignment, adaptive fusion, and task-specific supervision, M3A-DTA learns robust representations that accommodate incomplete modalities. The design enables joint training on sequence-centric data together with structure-centric complexes. We specify evaluation under complete, naturally missing, and deliberately masked modalities, with cold-drug, cold-target, and cold-pair splits. M3A-DTA achieves strong performance across multiple evaluation settings while demonstrating broader applicability than most existing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.