acceptodds
Under review as a conference paper at ICLR 2027

Coverage Is Not Selection: Exactly Attributing the Gains of Test-Time Scaling

Abstract

Test-time scaling methods (self-consistency, best-of- under a trained verifier, adaptive sampling) are compared by one number: accuracy at budget . Per question that number is an exact product: coverage , the probability a correct answer exists among samples, times selection, the probability of returning it given it exists; the first depends on the sampler, the second on sampler and rule. We turn it into an audit instrument: a remainder-free identity splitting any accuracy difference into base-rate, headroom and selection channels; an efficiency normalized between a content-blind floor and the oracle ceiling; and unbiased estimators pricing any rule at any budget from one pool of cached samples. On 149,000 generations from six models it does three things accuracy alone cannot. It attributes: a token cap and 63 hours at a fixed configuration leave accuracy unresolved at yet move the selection channel (cap: ; 63 hours: at , , and , , at ; one unreplicated pair). It prices verifiers: a 7B process reward model is the best selector we measure, yet selectors lose efficiency as the budget grows on all 9 (cell, selector) pairs (6 resolved, uncorrected), and how fast a vote loses it tracks how concentrated its per-sample weights are (Spearman with effective sample size, 17 pairs). It audits adaptive schedules at matched expected budget: adaptive-consistency gains on the three cells with ample coverage ( to ), though which channel pays is unresolved. All code, per-cell results and cached pools behind every number here are released at https://github.com/xxxxx/xxxxx (URL anonymized for review; public upon acceptance); code and results also accompany this submission, the pools being too large for it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.