Proving the Proxy Wrong: Witness Rank for Auditing Scientific Design
Abstract
Proxy-guided scientific design can rank far more candidates than trusted experiments can evaluate, yet standard metrics do not quantify the evidence for the single candidate ultimately selected. We ask how many trusted evaluations are needed to overturn that nomination. On a declared finite pool, the margin-aware *witness rank* is the position, under an ordering fixed without trusted labels, of the first candidate that improves on the nomination by more than . Across 230 audits on four fully measured protein and DNA landscapes, including two published offline optimizers, conditional median witness ranks are three to five despite the largest pool containing over 150,000 candidates. At sixteen trusted evaluations, proxy ordering achieves 1.68 to 8.60 times the exact random-order recall. Controlled rankings with Spearman correlation near 0.95 among the top candidates by measured performance still yield conditional median witness ranks of two to four, showing that high ranking quality alone does not determine audit cost. The same audit measurements identify replacements that cannot reduce measured performance, while independent random probes bound pool witness prevalence when no witness is found. To complement these retrospective biological audits, we demonstrate the auditing and replacement protocol in a prospective laboratory study of ionic conductivity in battery electrolytes. Together, these results make post-selection reliability measurable in experimental-budget units and guide whether to retain a nomination or replace it with a measured improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.