Oracle Coverage Can Reverse Rankings: Finite-Budget Identification in Test-Time Scaling
Abstract
Ranking systems by pass@K (oracle coverage) can disagree with ranking them by selected accuracy. We prove that strict oracle dominance at every sampling budget can coexist with plurality regret approaching one. We then derive sharp finite-budget identification bounds on plurality accuracy given correct-answer structure, abstention mass, and the largest wrong-answer probability. With these quantities fixed, extremal allocations of wrong-answer mass determine the accuracy range. Reanalysis of public response banks on 3,600 SuperGPQA problems finds high-budget ranking reversals between reasoning-effort settings after simultaneous correction, including when invalid responses participate in voting. On a 1,800-question development split at K = 64, adding the largest wrong-answer probability narrows empirical identification widths from 30.81–44.77 to 0.49–1.39 percentage points across three effort settings. These ranges describe information loss, not statistical confidence. We also give a fixed-score stopping certificate that skips verifier queries without changing the full weighted-vote decision. On 64 problems per policy, batch-preserving execution reduces verifier-stage time by 51–70%, excluding candidate generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.