acceptodds
Under review as a conference paper at ICLR 2027

N×N E-VALUATION: Hypothesis Certification Without Per-Hypothesis Conditional Resampling Models

Abstract

High-throughput discovery pipelines produce hypotheses faster than a conditional null can be designed for each, making certification a bottleneck. Held-out valida- tion shows whether a hypothesis predicts but not how: the same score can come from population prevalence, a pattern shared across the scope, a unit-matched re- striction, or an association carried by a minority. Conditional randomization tests must specify, for each hypothesis, its explanatory variables, conditioning set and resampling law. We introduce N×N e-valuation, in which the units governed by one hypothesis serve as references for one another. It tests whether held- out success is carried by the hypothesis’s scope, by its restriction, or by which unit receives which restriction, with no resampling model written per hypothesis, and reports which of these carries the prediction. E-values give finite-sample, per-hypothesis control of false certification under stated exchangeability nulls, and an asymptotically valid equivalence test certifies sharedness only when re- striction differences above a declared noise ceiling are rejected. On planted and randomized-intervention data, N×N certifies shared associations that a CRT and an exact permutation test reject; on these and on real data with expert-audited la- bels, it rejects minority-carried associations they certify: their nulls ask whether any unit-level dependence exists, not whether an association holds across its scope. Removing any of its statistics merges a verdict that certifies with one that rejects. The certificates concern predictive association, not causality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.