acceptodds
Under review as a conference paper at ICLR 2027

Commit Before Inspect: Machine-Enforced Outcome Blinding for Scientific Analysis Agents

Abstract

Scientific-analysis agents can favor a conclusion through analysis selection even when calculations are correct. Commit Before Inspect (CBI) withholds candidate results until an agent commits to an eligible plan, then binds evaluation to that choice. We test this auditable information boundary using matched visible and blind conditions with fixed candidate implementations. A negative pilot with nearly invariant selections motivates a diagnostic of the endpoint discordance required to observe an effect. In the prospective successor, three-model E1 estimates a 6.26-percentage-point reduction in false release (95% task-clustered interval, 4.03–8.50), below the prespecified 10-point meaningful-effect target, alongside a 6.56-point detection loss. Two-model E3 estimates a 3.01-point reduction (0.23–5.78), but conservative missingness analysis fails the registered significance threshold; detection falls by 5.00 points. The registered joint criterion is not met. In a separate 1,920-slot descriptive public-covariate supplement, matched visible minus CBI false release lies between 1.25 and 1.88 points under two missing endpoints; only five of 318 observed null pairs differ in endpoint. These results establish a testable information boundary, not a general scientific benefit or task-specific pursuit of favorable outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.