acceptodds
Under review as a conference paper at ICLR 2027

Controlling Property Trade-Offs in Reward-Alignment of 3D Molecular Generation

Abstract

Structure-based drug design requires balancing predicted binding affinity, synthetic accessibility, and molecular validity. However, models trained to reproduce the distribution of training ligands do not necessarily satisfy these objectives jointly. We adapt Flow-GRPO to a pretrained, pocket-conditioned 3D ligand generative model to jointly optimize multiple molecular rewards such as Vina docking and synthetic accessibility scores. Optimizing either objective in isolation improves the target metric at the expense of the other. By varying their respective reward weights, we identify the trade-offs between these conflicting objectives and relate them to the composition, topology, and flexibility of generated ligands. We find that, depending on the reward composition, the aligned model learns to balance these ligand descriptors to maximize the reward. This insight allows us to control the aligned models’ trade-off between Vina docking score and synthetic accessibility, with multiple configurations improving both over the baseline model, while also increasing drug-likeness and maintaining PoseBusters validity. Our results connect the trade-offs between objectives to interpretable changes in generated ligands’ chemistry, providing a basis for understanding and selecting useful compromises across protein targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.