Rethinking Directed-Evolution Benchmarks: Top-Variant Recovery Under a Fixed Budget
Abstract
Machine learning-guided directed evolution improves a protein over successive rounds of experimental screening. When variants must be assayed individually, as is common for catalytic activity, each round tests only tens of variants. Existing benchmarks such as ProteinGym reward ordering an entire mutational landscape rather than recovering the best variants under such a budget. We introduce **EvolveGym**, a benchmark of measured deep mutational scanning data that scores top-variant recovery under a fixed small screening budget, over three tasks drawn from practice: identifying beneficial single mutations, assembling higher-order multi-mutants, and searching a site-saturated library. It spans 487 single-substitution assays, 96 of them functional rather than stability assays, 44 multi-mutant landscapes, and 15 combinatorially complete enzyme-activity and binding landscapes. We further argue that variant selection is better posed as ranking variants against one another than as regressing their measured values, which are noisy and condition-dependent, and propose **EvolveRank**: it turns measurements into all pairwise comparisons and aggregates a learned preference classifier into a global ranking. Benchmarking 17 methods at a matched budget of 80 measurements, EvolveRank places 77.5% of its 16 nominations in the true top 5% of each single-mutation assay, against 41.0% for the strongest few-shot supervised baseline; with representation and budget held fixed, ranking adds nine points of precision over regression. On multi-mutant landscapes the winner is predictable from the landscape itself: additive models lead under near-additivity, EvolveRank under strong epistasis, and an ensemble of the two needs no prior knowledge of the regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.