acceptodds
Under review as a conference paper at ICLR 2027

UBOA: Unpacking Benchmark Outcomes and Assumptions for Properly Scoped Forecasting Claims

Abstract

A validation-selected forecasting rule may reject a 5% improvement null without supporting a guarantee conditional on its selection history. UBOA makes this distinction explicit through an executable specification for matching statistical claims to evidence. Each request binds the selected rule, comparator, loss, evaluation law, error event, conditioning information, and coverage unit. Eight rules either license the unchanged conclusion or retain an explicit mismatch alongside separately scoped evidence. Relative soundness requires correct annotations and external premises, which the checker does not authenticate. In a 140-case contract challenge fixed before access to the checker implementation, the revised matcher made no false accepts or false rejects; coordinate equality produced 29/70 false accepts and 10/70 false rejects. We also instantiate existing concentration tools in two synthetic settings, preserving raw mean absolute error (MAE) or mean squared error (MSE), overlapping horizons, and the selected rule while supplying finite-sample history-conditional local control. A preregistered audit retained 80 histories and 327,680 continuations, with one false promotion, zero recorded first-true-null inclusion violations, and no Holm-adjusted evidence of excess error. Supplementary coefficients and counts support deterministic reconstruction of the certificates and aggregate audit quantities. Earlier simulations and six real applications retain their historical interpretations; the real applications do not have these conditional certificates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.