PEPFM-TANRL: FINE-TUNING A MULTI-MANIFOLD FLOW MATCHING MODEL FOR PEPTIDE DESIGN VIA TANGENT-SPACE GRPO
Abstract
*De novo* design of peptides that bind a target protein is a discrete–continuous generative problem requiring the joint modeling of structure and sequence. Flow-matching-based generative models faithfully reproduce the distribution of native peptide–receptor complexes, yet likelihood training provides no guarantee that generated samples exhibit high binding affinity. We introduce PepFM-TanRL, a framework that brings Group Relative Policy Optimization (GRPO) to deterministic-ODE flow models operating over multi-modal continuous state spaces. Injecting time-decaying noise converts the ODE sampler into a stochastic policy, from which we derive factored importance ratios and analytic KL penalties across four modalities — , , torus, and logit space. As an architectural change that enables reinforcement learning, we predict torsion angles for all 20 amino acid types simultaneously and aggregate side-chain features via probability-weighted expectations, eliminating discrete sampling so that every generation step admits a closed-form Gaussian log-likelihood in the tangent space. Fine-tuning with a reward that splices the Vinardo docking score with a steric-clash penalty improves the reward distribution on all four targets (best-target median reward ; clash-free fraction 39% 80%); on the one target that starts with a very high clash rate, the gain stops at fewer clashes without reaching a Vinardo-scorable population. On the other three, Vinardo scores among clash-free samples also shift higher as energetically unfavorable conformations are suppressed, while native-like backbone geometry and structural diversity are preserved. We further analyze how reinforcement learning alters the similarity of generated structures to the native complex.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.