acceptodds
Under review as a conference paper at ICLR 2027

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

Abstract

SciCode evaluates whether language models can turn advanced scientific reasoning into working numerical code. Yet recent frontier models cluster near 60% subproblem accuracy, making progress difficult to distinguish. We show that this compression is partly caused by defects in the benchmark. We conduct a problem-by-problem, domain-expert audit of all 65 test problems and identify 264 defects. Of these, 192 reject correct instruction-following solutions and affect 91% of the main problems, through non-reproducible or incorrect gold values, over-tight tolerances, hidden numerical conventions, or contradictions between prompts and graders. Most score-suppressing defects require specialized physics or mathematics knowledge to detect. We correct every repairable defect, remove the one problem whose specification admits no verifiable target, and release the resulting SciCode-Verified benchmark with an itemized audit trail and reproducible grading harness. The corrections add the constraints required for well-posed tasks, repair grading, and strengthen tests that were too lenient; they do not lower the scientific difficulty of the problems. Re-evaluating twelve frontier models under a matched pass@1 protocol raises subproblem accuracy from 45–60% to 84–98% and main-problem accuracy from 9–27% to 69–92%. These results show that benchmark defects can compress both absolute scores and model rankings, and that domain-expert verification is necessary when evaluation targets scientific reasoning and code jointly. SciCode-Verified provides a more faithful and discriminating instrument for measuring scientific coding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.