Blip2Mol: Phenotype–Target Fusion for Molecular Generation via Multimodal Latent Diffusion
Abstract
De novo molecular design is commonly driven either by protein structure, which favors target affinity, or by phenotypic readouts, which better reflect system-level biological activity. However, existing methods rarely integrate both modalities in a unified generative framework, limiting their ability to produce molecules that are simultaneously bioactive and target-compatible. We propose Blip2Mol, a multimodal molecular generation framework that bridges chemical-induced transcriptomic responses, protein binding pockets, and molecular structures through a shared latent space. Blip2Mol first pretrains a molecular variational autoencoder to obtain a continuous chemical representation. It then aligns transcriptome and pocket embeddings to the molecular latent space via BLIP-style dual cross-modal alignment with contrastive and matching objectives, enabling semantic consistency across phenotype, target, and molecule modalities. Next, we develop a conditional latent diffusion model with a learnable gated fusion mechanism to generate molecules jointly conditioned on transcriptomic response and pocket context. Experiments on 10 cancer-related targets show that Blip2Mol consistently outperforms representative phenotype-guided and docking-based baselines, achieving stronger docking affinity while maintaining superior drug-likeness, synthetic accessibility, validity, uniqueness, and novelty. Ablation studies further confirm the effectiveness of multimodal conditioning and the critical role of the diffusion module. These results demonstrate the promise of our unified model for phenotype-aware and target-aware drug design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.