acceptodds
Under review as a conference paper at ICLR 2027

Leveraging Biological Priors for Protein Fitness Prediction

Abstract

Accurately predicting the properties of proteins is a foundational problem in protein science and engineering. Models trained across experimental assays can learn transferable patterns, but their success depends on the quantity and quality of available data. On the other hand, biological priors informed by signals such as sequence and structure are valuable for predictions in the low-data regime, but they may encode restrictive assumptions that do not accurately reflect empirical evidence. We introduce \method, a data-efficient protein fitness prediction framework that bridges these approaches. First, we incorporate sequence and structural information through protein language model embeddings and inverse folding features. Second, rather than relying solely on experimental data, we turn biological models into synthetic data generators, pretraining on the resulting fitness prediction tasks before finetuning on real assays. We find that synthetic pretraining improves downstream performance over training on experimental data alone, and our approach outperforms strong baselines across held-out ProteinGym assays.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.