acceptodds
Under review as a conference paper at ICLR 2027

Coverage Is Not Policy Value: An Audit of Test-Time Compute Oracles

Abstract

Test-time-compute studies often report an oracle that selects whichever stochastic action eventually succeeds on each problem. This measures retrospective coverage, which need not equal the value of choosing an action before its outcome is known. We distinguish retrospective coverage, the expected-action oracle (the best population action mean), finite-replication split-selected value, and held-out policy value. We show that coverage can grow with the number of exchangeable actions even when no decision-time policy improves over a fixed action. Our evaluation uses 771 model-problem states from three open models on MATH-500 and GPQA-Diamond, with five prompt actions and 16 outcomes per stochastic action. Under empirical product coupling, retrospective coverage exceeds split-selected value by 4.41 percentage points on MATH (95% problem-bootstrap interval [3.63, 5.26]) and 11.73 points on GPQA ([10.34, 13.17]); all six model-benchmark cells have positive gaps. Although split selection improves over a fixed prompt action, four router families evaluated on held-out problems show no statistically supported gain in either the prompt-action bank or a second bank that treats models as actions. Oracle headroom should therefore be reported with its outcome coupling and evaluated separately from decision-time policy value.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.