acceptodds
Under review as a conference paper at ICLR 2027

Proof-carrying hypothesis testing: certified statistical inference by AI

Abstract

Progress in formalisation technology, such as Lean, has allowed advanced mathematical proofs to be confirmed mechanically by computers. As well as providing greater assurance around existing, human, mathematical proofs, an application of this technology is to validate mathematical proofs by AI. In this paper, we consider the corresponding problem of certifying statistical inference by AI. We focus on testing a parameter (or regression contrast) under the linear model. We propose a statistical protocol, proof-carrying hypothesis testing, in which an AI produces a test with a formal proof of validity against the specific null hypothesis, design matrix, and assumption class set by the scientist, who withholds the data. An important aspect of the design of the protocol is to confront major tensions and statistical paradoxes which are pertinent to this hypothesis testing problem, namely the Bahadur–Savage theorem, Lindley's paradox, and the notion that “all models are wrong”. We demonstrate the protocol on thirteen datasets across science, policy, medicine and AI, with finite-sample p-values and e-values derived and machine-checked by AI, finding strong evidence in several cases under declared variance and error bounds. At those sample sizes, asymptotic p-values based on the CLT cannot be supported without further assumptions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.