acceptodds
Under review as a conference paper at ICLR 2027

Precision Inverse Folding: From Low-Order Mutant Measurements to Higher-Order Variant Prioritization

Abstract

Protein inverse folding models optimize sequence–structure compatibility but provide limited guidance on protein function. In practical protein engineering, only a small number of low-order mutants can typically be experimentally characterized, while promising designs may lie in a much larger space of higher-order combinations. We study this two-to-many setting, using only wild-type, single, and double-mutant wet-lab measurements to prioritize unseen variants with three or more mutations. Our approach constructs cross-order preferences from measured low-order variants and synthetic higher-order negatives derived from inactivity-associated mutations, focuses alignment on function-relevant mutations through critical-site preference optimization, and enriches structural representations with protein foundation-model embeddings. Across four protein fitness datasets and two inverse-folding architectures, our approach consistently improves higher-order variant prioritization, demonstrating that limited low-order experimental data can effectively guide exploration beyond the experimentally accessible mutational regime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.