BindGRPO: Training Protein Binder Generators to Bind
Abstract
Generative models of protein binders are typically trained by maximizing likelihood on natural complexes. Binding is considered primarily at inference: many candidates are sampled, scored by a structure predictor, and the highest-confidence ones are retained. Binding is therefore not part of the generator's training objective. We introduce BindGRPO to close this gap by post-training a target-conditioned binder generator using predicted binding as the reward. We formulate the denoising sampler as a Markov decision process and define Gaussian policy transitions between consecutive denoiser inputs with tractable likelihoods. The reward is AlphaFold3's interface confidence for the co-designed sequence. Training therefore does not require an experimentally determined target–binder complex. On the Cao benchmark, BindGRPO increases the in-silico success rate from to , a gain over the same generator before RL fine-tuning. The improvement persists under an independent AlphaFold2 evaluator and across most evaluation thresholds, suggesting that the gain is not specific to AlphaFold3. All code and data will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.