acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Check: Process Supervision with Artifact-Grounded Verification for Scientific Agents

Abstract

Language-model agents can search the literature, write analysis code, and operate scientific software. A prerequisite for improving such agents from experience is scalable feedback. Coding provides an instructive contrast: compilers and tests score attempts automatically and localize failures. In experimental science, decisive labels often arrive only after costly wet-lab work and reveal little about which intermediate decision failed. We ask whether the structured artifacts produced during scientific work can provide an analogous source of feedback. We focus on one capability within the broader scientific process: checking computational outputs before they guide experiments. Scientific pipelines often encode the same underlying quantities in multiple files, creating consistency relations that can be recomputed without access to experimental outcomes. We formalize review as three abilities (selecting what to check, executing the check, and interpreting its result) and construct a benchmark in which the relevant checks and their programmatic semantics are controlled. The resulting diagnosis is a check-elicitation gap: the agents we evaluate can execute a specified check but often fail to identify which check is relevant. This diagnosis motivates a model–verifier architecture. The model handles open-ended check selection and computation; missing check knowledge can be supplied at run time or distilled into its weights. An independent verification layer then recomputes claims governed by exact programmatic rules, without consulting task labels. Model-side interventions raise performance on withheld defects from 12.8% to above 88%, while verification restores accuracy on defect-free inputs to 91.7% without sacrificing those gains. The missing-check pattern also appears on an external biological protocol benchmark. More broadly, our results suggest a division of labor for artifact-producing scientific workflows: train flexible selection and computation into the model, and place machine-checkable claims behind an independent verification boundary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.