acceptodds
Under review as a conference paper at ICLR 2027

ScientistOne: Verifiable Autonomous Research via Chain-of-Evidence

Abstract

Autonomous research agents now produce competitive solutions and polished manuscripts, but their outputs can carry verifiability failures that presentation-focused evaluation misses: fabricated citations, unreproducible scores, and method descriptions that diverge from the code. These failures share a root: today's autonomous research systems are not built to trace their claims back to evidence. We address this with three contributions. First, Chain-of-Evidence (CoE), a verifiability framework under which every claim traces to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction across literature review, solution discovery, and paper writing. Third, CoE Integrity Audit, a post-hoc audit whose four integrity checks (score verification, specification violation, reference verification, and method–code alignment) apply the same criteria across systems. Across 75 papers from five systems on five frontier systems-research tasks, every baseline shows at least one systematic failure mode. Among the three baselines in the comparison, up to 24% of references are hallucinated, score verification passes in as few as 42% of papers, and method–code alignment ranges from 20% to 87%. ScientistOne is best or tied for best on all four checks, with none of its 340 distinct references hallucinated, 12/12 scores verified, and 14/15 method sections that match the code, while matching or exceeding the human expert solution on all five tasks. Against blinded human raters, the audit is right on all 48 score extractions and 219 of 220 passed references, and its method–code verdicts agree with the raters about as often as the raters agree with each other (92%). Our findings show that a deliberately simple discovery loop already reaches expert-level solutions, and we argue that the bottleneck in autonomous research is verifiability: CoE enables a paper to be audited before it is trusted, by both humans and machines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.