Molweaver: Scalable Foundation Model for Structure-based Drug Discovery
Abstract
Structure-based drug discovery requires designing molecules for target protein pockets and predicting ligand bound poses, both of which require learning molecular connectivity, three-dimensional structure, and ligand–protein interactions. We introduce Molweaver, a family of molecular foundation models ranging from 60M to 500M parameters, pretrained on 100 million SELFIES representations, 1 billion molecular conformers, 3.2 million protein pockets, and 167,000 redocked pocket–ligand complexes with associated properties. Molweaver is a unified transformer framework that learns molecular representations across 2D structure, 3D geometry, and chemical properties. It generates molecular sequences through autoregressive prediction and 3D conformations through flow matching. We adapt Molweaver to downstream structure-based tasks using supervised fine-tuning and reinforcement learning. For pose prediction, Molweaver places 80.2% of ligands within 2 Å RMSD of the reference pose while also passing all applicable PoseBusters checks, comparable to equivariant structure-based models despite using no equivariant inductive biases in the architecture. For molecule generation, docking-based feedback and molecular property optimization steer the model toward compounds with favorable interactions within CrossDock2020 target pockets, yielding substantially improved docking scores, full connectivity, and the lowest clash counts among compared methods. We also show that these metrics improve with model size. These results highlight the potential of large-scale joint pretraining and task-specific adaptation for developing general-purpose models for structure-based drug discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.