Self-Improving Bayesian Flow Networks via Group Relative Policy Optimization for Structure-Based Drug Design
Abstract
Structure-based drug design (SBDD) requires molecules that bind favorably to a protein pocket while maintaining plausible conformations and suitable chemical properties. Generative models learn from protein–ligand complexes, but fitting the training distribution alone does not directly optimize these properties. On-policy reinforcement learning offers a way for generators to learn from evaluations of their own candidates. Applying it to Bayesian flow networks (BFNs) requires identifying the stochastic decisions whose likelihoods support policy updates. In the molecular BFN sampler considered here, the network predicts clean coordinates and atom-type probabilities, while the sampler draws new coordinate-belief means and atom-type logits. We introduce GRAB (Group-Relative Alignment of Bayesian flow networks), a self-improving framework that defines these sampled quantities as a joint policy action. For this Gaussian sampler, we derive a closed-form conditional action density that is differentiable through the network predictions, with covariance fixed by the sampling schedules. Terminal rewards from docking affinity, strain energy, drug-likeness (QED), and synthetic accessibility (SA) support Group Relative Policy Optimization (GRPO), without a learned value model or gradients through reconstruction or scoring. Generation retains the original sampler without additional property-gradient guidance. On 100 CrossDocked2020 test pockets, one GRAB configuration achieves the lowest mean Vina Dock and median strain energy among the generators in the main comparison, while maintaining favorable drug-likeness, synthetic accessibility, and ligand diversity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.