Multi-Scenario Generative Protein Design via Preference Alignment
Abstract
Generative models can now design proteins that perform one function in one context: bind a given target, scaffold a given motif, catalyze a given reaction. However, many applications need a sequence that behaves differently depending on its context, like a binder that engages two unrelated targets, or a switch that adopts one fold on its own and another once a ligand is bound. Training a model for such multi-scenario tasks directly is impractical, since structural data of the same protein across contexts is scarce. To address this, we introduce Control-Proteina (Ctrl-P), which instead coordinates separately trained models at inference time, eliminating the need for training on multi-scenario data: each model expresses its preferences as a population of sequences and conditions its own generation on the preferences of the others, so that iterative refinement converges on designs that succeed in every context. In a multi-target protein–protein interaction benchmark that spans similar and dissimilar binding interfaces, Ctrl-P solves 18 of 20 target pairs, including 8 of the 10 challenging dissimilar pairs where the strongest baseline solves 2. We further demonstrate the generality of our approach by applying Ctrl-P to design allosteric motif switches that adopt specified conformations in response to a binding partner: it solves 8 of 12 tasks, compared to the strongest baseline’s 7, while finding more than twice as many successes per task on average. A case study on pan-arenavirus binder design shows that Ctrl-P generalizes beyond two scenarios, yielding candidates predicted to engage all three targets where no evaluated baseline yields any. Together, these results extend protein design toward programmable behavior across molecular contexts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.