From Tools to Instruments: Tracking Measurement Uncertainty in Autonomous Science Agents
Abstract
AI co-scientists increasingly do work that researchers have always done themselves: they test hypotheses on their own, from planning an analysis to interpreting its results. Yet inaccurate measurements from their tools can turn a correct analysis into a wrong conclusion. Such failures are hard to identify when reference answers are unavailable. We propose a framework that treats tools as measurement instruments and carries measurement assessments through autonomous hypothesis testing. Instrument disagreement supplies an assumed error scale, and a measurement counterfactual estimates the perturbation needed to change a statistical decision. Four checks combine measurement sensitivity and instrument disagreement with sampling uncertainty, verdict–evidence contradictions, and disagreement across repeats. Their flags define a risk level computed without reference answers. We implement the framework in VERITAS-M, a multi-agent clinical co-scientist, and evaluate it on hypotheses over cardiac and brain MRI, retinal images, and clinical data. VERITAS-M improves on its predecessor; using the same primary model with an additional critic, it also outperforms single-model agents in verdict accuracy. With each cohort's primary instrument, keeping conclusions flagged by at most one check lowers observed error about tenfold; the rule also lowers error for three systems it was not tuned on. Because inaccurate measurements can yield wrong conclusions that stay consistent across repeated runs, autonomous science needs what experimental science has long practiced: treating tools as instruments whose errors are measured and carried through to the conclusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.