acceptodds
Under review as a conference paper at ICLR 2027

From Chains to Trees: Efficient Exploration and Credit Assignment in Protein Diffusion

Abstract

Diffusion models provide strong generative priors for conditional protein design, but steering them toward downstream objectives remains challenging. Reinforcement learning (RL) post-training is crucial for aligning generative models with specific targets, but its prohibitive computational cost remains a major barrier to widespread adoption. In protein design, this leads to two problems: (1) rollout budgets are spent repeatedly resampling similar prefixes; (2) terminal rewards provide sparse supervision for long denoising trajectories. We propose TSG-Diff, an RL framework that organizes reverse diffusion as a branching tree over intermediate checkpoint nodes. Improves training efficiency by recasting the denoising process as a search tree. Starting from a small set of initial rollouts, TSG-Diff expands selected checkpoints to diverse continuations while reusing partial trajectories. Terminal rewards on final designs are converted into group-relative leaf advantages and then propagated through the tree to produce segment-level credit for intermediate denoising decisions. We instantiate TSG-Diff for protein design with a composite reward combining binding affinity, interface confidence, and structural regularization. Our approach provides more informative supervision under a fixed rollout budget, and empirical results show improved optimization of binding-related objectives while preserving structural plausibility and naturalness. The code can be found at this anonymous link: https://anonymous.4open.science/r/TSG-Diff-F223

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.