Controlling False Discoveries via Auditing Scientific Agents
Abstract
Scientific discovery agents generate and test hundreds of hypotheses per dataset, and the agent that explores also decides which findings to report. Its internal scores and per-hypothesis tests do not bound the fraction of false reports: in environments with known ground truth, 18-66% of agents' reported discoveries are false. We introduce audited discovery, which separates exploration from reporting. An unmodified agent explores one data split and registers its candidate findings, and an independent audit tests each candidate once on a sealed holdout and reports only those selected by a false discovery rate (FDR) procedure. Across five agent systems in synthetic and semi-synthetic (plasmode) environments, auditing the same candidate lists brings FDR below 1% for every agent, whereas fresh data or FDR correction alone leaves up to 51%. Reliability has a price. Four agents report 6-35% fewer true discoveries than in their own reports, while the fifth gains 63% from truths it had withheld. At a matched 5% FDR, auditing yields up to 26% more true discoveries than correcting exploration evidence, with the largest gains for agents that find the most truths. On real neuroscience data, blinded experts believed 57% of audited findings to be real, against 21% of findings that passed only an individual holdout test. Audited discovery lets agents scale their search while scientists fix the reported error rate in advance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.