Wrong Target or Failed Delivery? Diagnosing Scientific Coding Beyond Pass/Fail
Abstract
Strict scientific coding scores combine numerical computation, report construction, and final execution. A completion gap therefore need not represent an equally large difference in numerical content or program behavior. We examine this distinction through complete-reference matching, saved-content checks, and tests of frozen final programs linked to original execution records. In 43 historical starting-program pairs, strict completion differs by 65.1 percentage points, whereas the requested-vector match gap is bounded by 0–14.0 points. With reading rules fixed before model calls, GPT-5.4-mini and Qwen3.7 Plus show completion gaps of 66.7 and 75.0 percentage points, respectively, despite zero saved-content gaps. Haiku's completion gap exceeds its saved-content gap by at least 33.3 points; Gemini completes both starting conditions. A factorial design then varies computation and delivery support separately. Across six model configurations, the records distinguish correct numerical returns in invalid reports, numerical or return-type failures, and cases lacking qualifying final-source execution evidence. The prespecified Mini–Haiku replication yields an observed delivery-only versus computation-only completion difference of +20.8 points (). Joint evidence shows which completion differences accompany changes in tested computation and which coexist with matching numerical content and unfinished reporting or execution obligations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.