acceptodds
Under review as a conference paper at ICLR 2027

Anytime-Valid False Discovery Rate Control for Adaptive Foundation Model Evaluation

Abstract

Foundation model leaderboards are adaptive experiments: evaluators repeatedly choose model comparisons, allocate samples to promising claims, and stop when results look publishable, so standard fixed-time p-value procedures do not generally retain valid false-discovery guarantees after data-dependent stopping. We formulate leaderboard evaluation as sequential FDR control over directed pairwise model claims, and show that fixed-time BH fails under adaptive stopping while pairwise e-processes combined with e-BH control FDR at arbitrary stopping times. The main practical obstacle is power: greedy allocation concentrates nearly all budget on one comparison, leaving other true alternatives undiscovered. We make this quantitative: greedy allocation exhibits a 1/K power ceiling up to an explicit finite-sample correction (K true alternatives), allocation-aware sampling recovers power, and we prove a Theta(m/s) separation in discovery budget between top-k and uniform sampling that holds for any FDR-controlling rule (not only e-BH). We propose allocation-aware anytime-valid protocols (top-k round-robin, threshold reallocation) and validate them on synthetic adaptive leaderboards and adaptive replays of Chatbot Arena and MT-Bench, showing inflated fixed-time error rates, stable e-BH control, and substantial power recovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.