ShrinkFlow: Learning to Shrink Proteins from Protein Language Model Feedback
Abstract
Protein shrinking seeks shorter variants of existing proteins, yet experimentally validated source–shortened pairs are scarce and substantial shortening may require coordinated deletions across multiple regions. We introduce SHRINKFLOW, a framework for learning reusable multi-region shrinking policies from protein language model (PLM) feedback. A frozen PLM evaluates complete shortened candidates relative to their source, and the resulting plan preferences are converted into history–action supervision for a conditional discrete interval flow. A feasible action space enforces exact deletion budgets while preserving the identity and order of retained residues. Across general proteins and protein complexes, SHRINKFLOW achieves improved predicted structural and interface preservation relative to recent learning-based shrinking methods. Under shared functional-site constraints, it also preserves the local structural environment around protected catalytic regions more effectively. Controlled studies show that complete-outcome evaluation improves candidate selection over WT-only scoring and that the released policy achieves higher first-output structural quality than a capability-stage reference. Together, these results support reusable, exact-budget protein shrinking from sequence-model feedback without experimentally paired shortened targets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.