ProEdit: One-Step Protein Optimization with Diffusion Protein Language Models
Abstract
Protein design often begins downstream optimization from a complete candidate sequence, yet reward-guided optimization of diffusion-pretrained protein models typically retains the native iterative denoising process. We ask whether this generative trajectory is necessary once a candidate already exists. We introduce ProEdit, which repurposes a diffusion-pretrained protein language model as a target-specific protein sequence optimizer. Rather than optimizing through the native denoising process, the editor conditions on the complete parent sequence and samples a complete edited sequence as a single parallel action. This policy admits explicit action likelihoods and an exact conditional KL to a frozen reference policy, enabling direct reward learning without reconstructing diffusion trajectories. We train the editor from black-box structural rewards using parent-wise group-relative reinforcement learning, where candidates sampled from the same parent provide the relative optimization signal. On protein-binder optimization against PD-L1 and IFNA2, ProEdit substantially outperforms native diffusion-rollout GRPO in reward and iPTM while reducing training GPU-hours by by a factor of approximately 27 under matched reward-evaluation and optimizer-update budgets, and also outperforms reward-weighted supervised training. Most of the improvement is already present in raw policy proposals before reward-based selection, and the gains persist under independent structure predictors. These results show that diffusion pretraining can provide a strong initialization for downstream protein optimization without prescribing the decision process used to optimize complete candidate sequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.