MGPO: Mutant Group Policy Optimization for StructureDiverse Protein Reinforcement Learning
Abstract
Proteins can admit multiple structurally distinct solutions for the same amino-acid sequence, making broad structural exploration important for protein generative models. Mutation offers a natural way to expand policy exploration beyond ordinary autoregressive sampling, but we identify a previously under-characterized failure mode of mutation-augmented reinforcement learning: evolutionary exploration produces heterogeneous structural neighborhoods with different reward offsets, so a single global baseline can assign negative advantages to mutants that are improving within their own neighborhood. We introduce Mutant Group Policy Optimization (MGPO), which couples structured evolutionary exploration with neighborhood-aware credit assignment in the protein-structure-token space. MGPO generates structurally diverse mutant populations through reward-blind mutation and quality-diversity search, organizes them using reward-independent structural geometry, and computes each mutant's advantage relative to its full structural neighborhood before representative compression. This preserves learning signals for promising mutant families that would otherwise be suppressed by global population statistics. On the CAMEO22 benchmark, MGPO substantially improves mean lDDT-CA and TM-score. Diversity is evaluated in both structure token space and decoded 3D space; the measurements show that MGPO can improve the diversity of generated protein structures. A matched same-layout comparison, together with local-only ablations, shows that neighborhood-local credit assignment accounts for most of the improvement, while the evolutionary search provides the diverse structural neighborhoods on which this mechanism operates. These results establish MGPO as a mutation-aware policy-optimization framework for jointly improving structural quality and maintaining population-level exploration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.