Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science
Abstract
When an agentic science system improves against a fallible evaluator, what establishes a new capability rather than evaluator-specific optimization? We propose a two-sided audit that separates an exact exclusion for the prior verifier from empirical gains under a shared adjudicator. RNA topology supplies the negative side: a pseudoknot-free verifier cannot certify crossing structures at any search budget. The positive side is model-relative, assessed by three structure predictors with one held outside every optimization loop. The audit returns six findings rather than a pass/fail bit. A hand-built operator solves 43 of 60 crossing targets under the predictor it optimizes, but only one under all three. On the same 43 targets, an unoptimized predictor confirms two of its designs versus 26 for an external solver. Auditing our revised operator finds a narrow improvement but no measured advantage over matched-cap undirected search. The instrument nevertheless discriminates: two failure-free frozen LLM-written operators attain held-out carry-over of 0.293 versus 0.095 on 951 paired units, using 4.6–10 times fewer oracle calls. Seven interventional tests do not explain the difference. These results support externally adjudicated, compute-accounted capability comparisons, not physical validation, a policy-level grammar transition, or discovery of a transferable design principle.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.