acceptodds
Under review as a conference paper at ICLR 2027

Replay the Scientist: Calibrated Evidence After Interactive LLM Hypothesis Search

Abstract

Large language models increasingly search for scientific claims through interaction. They adapt hypotheses, analyses, repairs, and measurements before reporting a result, but testing only the final hypothesis discards the outcome-dependent path that selected it. We introduce Transcript Replay, a black-box protocol that reruns the discovery procedure under randomized outcomes and ranks the selected score against counterfactual discoveries. Under a declared exchangeable null, the discovery program is the inferential unit and retains its search policy. In controlled studies, final-only rejection reaches 96.3%, while Transcript Replay remains calibrated through adaptive search and acquisition. We evaluated Qwen and Llama on rule search, generated Python features, and external equation discovery. In full-data rule search, Transcript Replay has 3.0% null rejection versus 64.8% for final-only testing and gains 24.0 power points over held-out testing. At fixed full-data scoring, a larger search sample improves power by 11.5 points and planted-mask recovery by 12.2 points. We reuse exact reference scores, reducing required executions by 91.9% in the primary validation batch for repeated queries to a fixed scientist on outcomes related by the specified transformations. The replay *p*-value evaluates the discovery procedure behind the selected rule or program.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.