Joint Foundation Model for Synthesizable Molecular Generation and Property Prediction
Abstract
Molecular foundation models deliver strong property prediction from limited labels but produce encoders that cannot propose new molecules, while latent-variable generators are trained for reconstruction rather than transfer and decode unreliably away from their data. Design needs a third ingredient neither supplies: assurance that what is proposed can actually be synthesized. The model presented here, that is a generative version of the foundation model CheMeleon, puts all three on one representation. It pairs the descriptor-supervised CheMeleon D-MPNN with an autoregressive transformer decoder, optimizing Mordred descriptor regression and SMILES reconstruction jointly at foundation scale, so that a single -dimensional embedding serves as property representation, generative latent and search domain. A rectified flow fit to cached embeddings replaces the intractable Gaussian prior and doubles as a projection operator that keeps latent search decodable. Synthesizability enters the same space rather than a post hoc filter, as a differentiable head distilled from a retrosynthesis planner and, in a second arm, as hard projection onto routed analogs. Invertibility costs RMSE on property prediction across 30 MoleculeACE assays. The generative capabilities, tested on the GuacaMol benchmark, are competitive to state of the art models for both unconditional and goal-oriented generation tasks. Driven by property-head gradients under Langevin dynamics, the model proposes novel, potent, synthesizable candidates on MoleculeACE assays, each with an explicit synthetic route.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.