What Did the Verifier Verify? Auditing Evidence Attribution in Scientific Agents
Abstract
Scientific agents can reproduce numerical results while assigning them to unsupported claims. We audit their verification pipelines by separating execution-evidence representation, original-claim representation, and deterministic checking. Historical workflows motivate a controlled follow-up with 24 assertions arranged in 12 workflow-level pairs, two model configurations, three review interfaces, and two repetitions. All interfaces can access the same execution artifacts. Luna's direct review makes one false authorization per repetition, whereas full structured review makes none before deterministic checking. GPT-5.5 direct review is correct on every item in both repetitions. Ordinary and specialized implementations of the same checking contract agree on all 192 structured decisions. Replacing extracted evidence with executor records repairs one workflow-level error in one compact-extraction run, without a repeated advantage over full structured review. The follow-up does not reproduce the extraction-induced degradation observed in exploratory cases. Our component-audit framing is a retrospective synthesis of these results, not confirmation of that hypothesis. It makes performance attribution auditable through matched ordinary controls, intermediate judgments, and field-level traces. In this controlled suite, the observed safety improvement is already present upstream of deterministic checking: evaluating only final verdicts would obscure both its location and remaining representation errors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.