Beyond Execution: Operationality and Scientific Fidelity Are Distinct Capabilities of Scientific Coding Agents
Abstract
Scientific coding agents are increasingly judged by whether the artifacts they produce execute. For simulation code, execution is only operational evidence: a case can run while instantiating the wrong physical parameter, geometry, boundary condition, or numerical intent. We ask whether *operationality* and *scientific fidelity* are the same capability. We encode five canonical OpenFOAM problems as hidden machine-readable scientific contracts and evaluate six contemporary language models under real solver execution and deterministic requirement checks. Both mismatch directions occur: executable artifacts violate the intended science, and scientifically faithful artifacts fail to run. Across the six evaluated models, the two rankings have a Spearman correlation of , and model ordering can reverse between the criteria; the reversal between the best executor and the most faithful model persists across three independent generations. We then intervene on the information the agent receives: given only ordinary runtime feedback, operationality rises by while scientific fidelity changes by , and 79% of doubly-failing artifacts whose execution is repaired remain scientifically wrong. Execution feedback repairs operationality but provides little signal for hidden scientific errors. Evaluation that collapses the two axes into a single execution signal can change which model is judged best.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.