Proteina-Reforza: Reinforcement Learning Protein Binder Design
Abstract
Flow- and diffusion-based generative models have transformed de novo protein binder design, but they are typically trained to reproduce molecular data distributions without explicit alignment to design objectives. Here, we introduce Proteina-Reforza, a reinforcement learning (RL) approach that post-trains Proteina-Complexa to improve protein binder design. We adapt the Advantage Weighted Matching RL framework to Proteina-Complexa’s partially latent representation for atomistic joint sequence–structure generation. Rewards include structure prediction confidence scores, physics-based Rosetta energies, affinity predictions, as well as a novel diversity reward, and we additionally train reward models on experimental binding data to enable RL guided by experimental outcomes. We apply Proteina-Reforza to both minibinder and cyclic peptide design, for both protein and ligand targets, and demonstrate general model alignment as well as target-specific adaptation for challenging design tasks. Proteina-Reforza achieves state-of-the-art performance on established in silico binder design benchmarks, substantially outperforming existing methods, and alleviating the need for extensive post-hoc filtering. Further analyses and ablations reveal how RL shapes the generated binders and support their structural and biophysical plausibility, without noticeable reward hacking. Experimental evaluation of generated minibinders and cyclic peptides confirms that these computational gains translate into improved binding success in the lab.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.