acceptodds
Under review as a conference paper at ICLR 2027

Scoring How, Not Whether: A Protocol for Evaluating Discovery in Scientific Agents

Abstract

Scientific research combines prior knowledge with evidence, and a discovery is usually made where the prior is absent or wrong. An agent can therefore contribute to discovery only if it derives claims from the evidence and lets that evidence override what it already knows or has been told. A correct answer alone cannot show which happened: reasoning from the evidence, recall from a prior, or acceptance of an answer it was handed. We therefore propose a protocol that evaluates the process separately from the outcome, combining controlled disclosure of the target identity, a planted false identity, and three trace measures: where each claim came from, whether it is backed by computed evidence, and whether a supplied identity was tested against the agent’s own. The measures are read from the trace by judge models from outside the evaluated families, so that no model grades its own, and each is paired with a control condition or an input ablation. We validate the protocol on cancer subtype discovery with GPT-5.5, Sonnet 5, Opus 5 and Gemini 2.5 Pro. Four findings emerge. Withholding the cancer identity does not keep it hidden: GPT-5.5 and Sonnet 5 infer it from the gene symbols, Opus 5 names it before any gene is visible, most often from the sample count and clinical profile, and Gemini 2.5 Pro most often names none; the agent then confirms the cancer rather than considering alternatives, and the outcome score records none of this. The agents are better interpreters than analysts: all four propose well-supported mechanisms, yet once the identity is withheld three recover the published subtypes no better than a scripted k-means, and only Opus 5 beats it in every condition. A prescribed procedure improves the outcome score with no detectable reduction in adoption of a planted false identity. When a supplied identity conflicts with the agent’s own evidence, Sonnet 5 reads its evidence as supporting the label in every episode it tests, GPT-5.5 records the contradiction and defers, Gemini rarely tests the label, and Opus 5 follows its evidence while accepting every true label; what separates them is which side prevails when a check fails. Human annotators agree with the judges on which side prevailed, less so on whether the label was tested. The results indicate that current agents are not yet reliable enough to be left unsupervised on a discovery problem: their answers can be recalled or accepted rather than derived, and a correct one does not show which. The protocol provides the means to examine this from the agent’s log, and can be applied where no answer is available to check against.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.