acceptodds
Under review as a conference paper at ICLR 2027

ProteinFitnessBench: Learning to Extrapolate Protein Fitness Across Mutation Orders

Abstract

Protein engineering often requires selecting a small number of higher-order variants from models trained mainly on wild-type and low-order mutants. We introduce ProteinFitnessBench, a controlled framework that separates representation, readout, objective, and evaluation, and reports fixed-budget true-top-K recovery (TTK@K) alongside global ranking metrics. Across five seeds and four protein-fitness landscapes, a tested ΔESM-WHT readout package improves global ranking relative to a matched ΔESM-MLP comparator, but does not yield a consistent TTK@10 gain. This exposes a key evaluation mismatch: better global ranking does not necessarily produce better elite retrieval under a finite assay budget. Supporting GB1 analyses under sparse supervision and altered label resolution reinforce this distinction. ProteinFitnessBench provides a controlled way to test whether improvements in protein-fitness models translate into experimentally useful candidate selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.