acceptodds
Under review as a conference paper at ICLR 2027

AI Scientists Need Stronger Verification Framework to Address Systemic Fragility

Abstract

Over the past few years, the development of AI Scientists has accelerated rapidly, demonstrating growing potential for automatic research across multiple disciplines. Despite this progress, the capability boundaries of current AI Scientist systems remain insufficiently characterized, especially with respect to verification ability. In this paper, we argue that autonomous verification should be treated as a first-class challenge in the development and evaluation of AI Scientist systems. Current AI Scientist systems exhibit clear verification shortcomings, demanding more comprehensive strategies to address the systemic fragility of autonomous verification. We support our position with three levels of evidence. First, we review coarse-grained evidence from existing benchmarks on complex scientific tasks, which shows that frontier models remain unreliable when tasks require domain-specific reasoning. Second, we conduct an output-level evaluation of 28 research papers generated by five AI Scientist systems, revealing persistent weaknesses in experimental design and methodological clarity. Third, we introduce DeepVerify-5K, which contains 4,890 end-to-end scientific discovery trajectories. Each trajectory provides records of idea generation, implementation, execution, and iterative revision, enabling process-level analysis of how autonomous verification fails. Our analysis of DeepVerify-5K shows that current AI Scientist systems often exhibit fragile verification behaviors, including redundant code changes and entanglement with evaluation pipelines. We hope this position paper encourages the community to pay greater attention to the systemic fragility of autonomous verification and evaluate AI Scientist systems' capabilities in closer alignment with actual scientific scenarios.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.