acceptodds
Under review as a conference paper at ICLR 2027

SCIRIGOR: EVALUATING OPEN-ENDED SCIENTIFIC ANALYSIS BEYOND FINAL SCORES

Abstract

Scientific coding agents produce code, results, figures, and written findings within a single run. A correct final claim can nevertheless accompany contradictory anal- ysis or incomplete evidence. We introduce SCIRIGOR, a benchmark that evalu- ates both claim correctness and support from the run’s computations, figures, and analysis. It contains 100 cases from 27 articles across six domains and 17 sub- fields, with source code, inputs, results, and figures for verification. An evidence graph connects each claim to numerical results, plotted comparisons, and written interpretations. The evaluator checks these records against reference requirements and tests whether they support one another. Each claim’s support is limited by its weakest required check, while approved alternatives accommodate valid methods and visual encodings. Across eight model–runner configurations, the evaluator identifies incomplete support in 64 of 357 submissions with perfect final-claim scores. The highest Path F1, which measures supported claims and finding cover- age, is 84.4%; no configuration exceeds 40.8% Strict success, which requires all mandatory checks to pass.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.