MOLFRA: Online Alignment of Multimodal Molecular Flows via Forward-Process Reinforcement Learning
Abstract
Pretrained generative models for structure-based drug design (SBDD) learn structural regularities from protein–ligand complexes without directly optimizing task-specific molecular properties. Controllable approaches pursue these objectives through sampling-time guidance or iterative optimization, but improving properties while retaining molecular diversity remains a challenge. We introduce **MOLFRA**, an online forward-process reward-alignment method that jointly fine-tunes pretrained multimodal molecular flows. Under a shared pocket-relative reward, MOLFRA couples implicit positive and negative prediction branches in coordinate endpoint and atom-type logit spaces, entirely avoiding sampling-trajectory likelihood ratios or backpropagation through the sampler. On CrossDocked2020, post-training MOLFORM with a joint Vina Score and synthetic accessibility (SA) reward significantly improves mean Vina Dock (from −7.50 to −9.25 kcal/mol) and normalized SA (from 0.60 to 0.74), while fully preserving the pretrained baseline’s molecular diversity (0.78 vs. 0.78).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.