Coverage Is Not Policy Value: An Audit of Test-Time Compute Oracles
Abstract
Test-time-compute studies often report an oracle that selects whichever stochastic action eventually succeeds on each problem. This measures retrospective coverage, which need not equal the value of choosing an action before its outcome is known. We distinguish retrospective coverage, the expected-action oracle (the best population action mean), finite-replication split-selected value, and held-out policy value. We show that coverage can grow with the number of exchangeable actions even when no decision-time policy improves over a fixed action. Our evaluation uses 771 model-problem states from three open models on MATH-500 and GPQA-Diamond, with five prompt actions and 16 outcomes per stochastic action. Under empirical product coupling, retrospective coverage exceeds split-selected value by 4.41 percentage points on MATH (95% problem-bootstrap interval [3.63, 5.26]) and 11.73 points on GPQA ([10.34, 13.17]); all six model-benchmark cells have positive gaps. Although split selection improves over a fixed prompt action, four router families evaluated on held-out problems show no statistically supported gain in either the prompt-action bank or a second bank that treats models as actions. Oracle headroom should therefore be reported with its outcome coupling and evaluated separately from decision-time policy value.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.