Winners Need Witnesses: Certified Replay Selection and Reporting with PAIRS
Abstract
Selecting a good configuration and estimating its performance are distinct statistical tasks. We formulate joint selection and reporting: under a compute budget, a procedure returns an ε-optimal configuration and an η-accurate performance estimate with a joint certificate, or abstains. Shared randomness can make comparisons nearly noiseless while absolute-performance estimation remains costly. We establish a finite-pool minimax characterization and, for two candidates with known Gaussian covariance and complete-pair queries, instance-dependent bounds matching up to logarithmic factors. PAIRS combines paired sequential selection with protected reporting. A report-aware controller has a variance- and gap-dependent completion bound; paid variance learning can change the cheapest acceptable winner. Under the stated information-use assumptions, reports are selection-conditionally unbiased and incorrect joint certification has probability at most δ. Four replay benchmarks yield reference-relative Pε = 93.2%, versus 81.1% for uniform allocation at equal mean selection cost. In a 20-run Split-CIFAR-10 task-order study, fresh same-winner reporting reduces referencerelative macro absolute signed bias by 92%. A 1,000-pair application evaluation gives a 2.30 ratio between the smallest evaluated budgets meeting prespecified completion and correct-joint-output targets, with reporting and audit charged and abstentions counted as failures. A heterogeneous-variance study shows that additional comparisons can reduce total cost by shifting selection toward an easierto-report acceptable candidate. PAIRS makes reporting precision and completion explicit objectives of replay selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.