ProteinREPA: Structure Predictors as Representation Teachers for Generative Protein Design
Abstract
Modern protein design commonly follows a design-and-filter strategy, in which generative models propose candidates and biomolecular structure predictors determine which designs advance. In this paradigm, the predictor’s knowledge is primarily used to reject unsuccessful designs rather than to improve the generative process itself. We ask whether internal representations encoding geometric and evolutionary priors can be instead learned by generative protein design models and, if so, which capabilities improve. We introduce ProteinREPA, which aligns the pair representation of RFD3 with that of a frozen, MSA-conditioned Boltz-2 teacher using a lightweight learned projection and cosine-similarity objective. The teacher is used only during training and introduces no additional cost during standard inference. We compare ProteinREPA with matched unaligned controls trained using the same data and optimization schedule, evaluating unconditional sequence–structure generation, protein–protein binder design, protein–ligand binder design, and enzyme motif scaffolding. ProteinREPA improves intrinsic sequence–structure compatibility: the fraction of jointly generated sequence–structure pairs that refold self-consistently rises from 34.5% for the matched control to 43.8% with the affine projection and 51.3% with the bias-free linear projection. Furthermore, training RFD3 with ProteinREPA improves the success rates for the majority of task specific design benchmark cases, such as enzyme–motif scaffolding, protein–protein-, and protein–ligand-binding tasks, demonstrating that alignment also improves functional scaffold generation. Increasing coordinate-mediated self-conditioning from two to ten iterations provides a further and often larger gain across the three benchmark classes; for example, the number of passing BHRF1 binder backbones increases from 63 to 125 out of 200. Together, these results establish biomolecular structure predictors as representation teachers for generative protein models and show that improved learned representations and inference-time geometric search address complementary limitations of current all-atom generators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.