acceptodds
Under review as a conference paper at ICLR 2027

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

Abstract

LLM agents are increasingly used to reproduce and extend scientific experiments, and their output is often judged by whether the generated code runs and reports plausible metrics. We argue that this criterion is insufficient and study methodological hallucinations: silent deviations from a reference experimental protocol that leave execution intact but undermine the scientific conclusion. Examples include reducing datasets or training budgets without reporting it, replacing a failed learning or generative component with a lookup or oracle function, and drawing conclusions at a scale where the method's claimed advantage cannot appear. From a case study of 30 long-horizon reproductions of machine learning papers under a single consumer-GPU budget, we derive a five-class taxonomy of such failures. 17 of the 30 runs (56.7%) exhibit at least one, even though every run was produced by an agent that had the reference protocol in its prompt. To make such deviations easier to surface, we describe ABE-Ralph, a lightweight auditing protocol that encodes a paper's claims, required components, baselines, and metrics as a YAML contract, guides implementation through an 8-step workflow, and applies three inexpensive post-hoc checks: result coverage, LLM-based semantic review, and module presence. In a single-run comparison with four agent baselines, ABE-Ralph obtains the highest composite score (58.8, versus 51.0 for the strongest baseline) on a 21-task set whose per-system coverage differed (10–18 of 21 tasks). Because the systems also differ in backbone model and iteration budget, however, we cannot attribute this gap to auditing alone. Our findings suggest that evaluations of AI scientists should check whether an experiment faithfully tests the intended claim, not only whether it runs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.