Online Search-Augmented Reinforcement Post-Training for All-Atom Protein Design
Abstract
Improving pretrained all-atom protein generators toward downstream biophysical objectives remains challenging but is underexplored. Reinforcement-based post-training can directly adapt the generator, but trajectory-level optimization over hundreds of denoising steps is expensive. Inference-time search can instead improve candidate quality through additional sampling and selection, but leaves the generator itself unchanged and must be repeated whenever new candidates are generated. We introduce Search-Augmented Reinforcement Post-Training (SARP), which combines trajectory-free endpoint-based reinforcement with temporary supervision on high-reward endpoints discovered by reward-guided search. Across small-molecule binder and enzyme-design tasks, SARP improves design quality under standard sampling of the partially latent protein all-atom generator. In small-molecule binder design, SARP decreases binding energy by up to 14.6 REU and RF3 min-iPAE by up to 2.2 Å relative to the pretrained model. Across five enzyme targets, SARP produces 42.2 unique successes per 100 designs, compared with 11.2 for the pretrained model and 35.6 for the RL-only method DGPO. These gains persist under independent AlphaFold3 cross-evaluation, suggesting that the improvements transfer across structure predictors rather than arising solely from optimization of RF3-specific signals. Gradient analysis further shows that search supervision increasingly conflicts with reinforcement over training and that the backbone and latent flows experience different optimization pressures, suggesting that adaptive control of supervision and regularization may be beneficial. Overall, SARP provides a practical post-training framework for protein generative models that enables fast standard-sampling generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.