Can Scientific Agents Support Their Causal Claims with Experiments?
Abstract
Can scientific agents autonomously complete experiments that support their own causal claims, even when their models are correct? We introduce WORLDHATCH, a diagnostic benchmark whose main design holds a two-path structure fixed. Agents retain, revise, or expand a documented explanation. The design separates evidence that an influence exists, acts selectively, and disappears under isolation. We assess model correctness and successful operation separately from these measured comparisons. In the main evaluation of seven language models, the strongest agent correctly models and uses the additional influence in every episode where it exists, but completes all three evidence requirements in only 0.5%. We characterize these evidence gaps with executable alternatives allowed by the task that fit the recorded measurements while contradicting a causal claim in the submitted model. Across models in the main evaluation, verified continuations complete the missing evidence and another successful operation within budget for almost all otherwise successful but evidence-incomplete episodes with that influence, without changing their models or policies. In a matched prompt comparison, specifying all three verification targets without prescribing an experiment sequence yields near-complete evidence collection for one frontier model but leaves large gaps for two others. The correctness–evidence gap also appears in an asymmetric cross-path extension. Evaluation should distinguish correct modeling, guided verification, and autonomous evidence completion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.