From Plausibility to Reliability: Failure-Aware Closed-Loop Learning for Protein Design
Abstract
Protein design is increasingly able to produce large sets of computationally plausible candidates, while physical measurements remain the scarce resource that establishes reliability. We study folding stability as a controlled reliability endpoint and formulate measurement selection as a sequential learning problem under sparse physical feedback. We test whether an acquisition objective that combines model disagreement with predicted low stability can improve discovery of low-stability candidates under a fixed budget. PRIB Failure-Aware implements this objective as a scale-free rank aggregation of bootstrap-ensemble disagreement and failure relevance. In a frozen retrospective replay over 30,000 MGnify Stability sequences, each acquisition round reveals a batch of previously hidden experiment-derived targets and refits a three-member Ridge ensemble on frozen ESM-2 representations. Across five algorithmic seeds and cumulative budgets from 512 to 8,192 labels, PRIB achieves higher low-stability recall than Random and generic uncertainty sampling at every checkpoint. At 8,192 labels, recall is , compared with for Random and for Uncertainty; at 4,096 labels, PRIB already exceeds both baselines at 8,192 labels. On this fixed pool, the direction holds in all 25 seed-by-budget comparisons against each baseline. A target-provenance audit fixes the released convention before production replay, and predictive test evaluation is reserved for the final 8,192-label checkpoint. These results show that the acquisition objective can materially change which low-stability risks are resolved within a finite measurement budget and motivate experimental allocation as a first-class component of protein reliability workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.