acceptodds
Under review as a conference paper at ICLR 2027

LAVEP:LANGUAGE GUIDED PROTEIN LANGUAGE MODELS FOR VARIANT EFFECT PREDICTION

Abstract

Predicting how sequence variants alter protein function is central to protein engineering and variant interpretation. Yet protein fitness is assay-dependent, while most zero-shot protein predictors are assay-agnostic. Protein language models (PLMs) provide transferable sequence priors, but typically assign a fixed score to a mutation regardless of whether an experiment measures binding, activity, stability, expression, or another phenotype. Supervised methods can capture assay specificity, but rely on target-assay labels that are experimentally costly and time-consuming to obtain and often scarce in practice, limiting their use for new assays. We introduce a natural-language-guided framework that uses assay descriptions to tell a frozen PLM what “fitness” means in the experiment at hand. We treat language as an open-vocabulary task specification: assay context is grounded in the protein representation and used to steer the PLM's mutation prior toward the measured phenotype, without fine-tuning the PLM or requiring target-assay labels. We further establish a fundamental limit of assay-agnostic prediction. When the same variants are ranked differently across assays, any shared predictor must compromise between distinct fitness landscapes, imposing an intrinsic ceiling on attainable performance. Conditioning on assay language allows the model to represent these experimental objectives separately and escape this shared-ranking constraint. Across homology-separated ProteinGym assays, our approach consistently improves zero-shot mutation-effect prediction on unseen assays across ranking and classification objectives, showing the value of assay language precisely where new experimental labels are limited.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.