acceptodds
Under review as a conference paper at ICLR 2027

Euterpe: Efficient Multi-Level Preservation of Pretrained Performance during Preference Alignment for Protein Inverse Folding

Abstract

Protein inverse folding (IF) aims to design amino acid sequences for a given protein backbone, making it a fundamental tool for protein design and engineering. Recent Direct Preference Optimization (DPO)-based IF post-training improves developability properties such as solubility and thermostability, but can drive the aligned model away from the pretrained IF model. Existing multi-objective methods address this trade-off by incorporating costly folding-derived structural signals as additional designability alignment objectives. Our analysis shows, however, that the resulting structural gains largely recover the performance lost during developability alignment, rather than improving beyond the pretrained IF model. This raises a natural question: can we regulate developability alignment using the pretrained IF model itself, without costly folding-derived feedback? We introduce Euterpe, an efficient semi-online preference optimization framework that uses the pretrained IF model as a reference prior to regulate alignment. Pair-level Reweighting uses reference-model energies to downweight preference pairs weakly supported by the pretrained distribution, while Token-Level Thinking-Pattern-Guided Reweighting modulates residue-wise updates according to the thinking pattern induced by the pretrained IF model. Built on ProteinMPNN, Euterpe preserves designability and AAR at levels comparable to methods using explicit folding-derived structural feedback, while eliminating such feedback during alignment, requiring fewer GFLOPs, and achieving up to a speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.