The Hidden Allowance Axis in Test-Time Compute Comparisons
Abstract
A parallel test-time compute baseline chooses both sample count and per-sample token allowance. We show that fixing the allowance can conceal stronger baselines and change strategy comparisons. Two studies each select configurations on 200 validation questions and evaluate them on 400 held-out questions with three new test seeds. On MATH level 5 with two models, freeing the allowance at fixed sample count raises pooled accuracy by 7.58 points, at 23.4% higher input-plus-output token cost. A preregistered extension to four non-mathematical MMLU-Pro categories and three models opens the same allowance/count search to natural and reconsideration-suffix sampling. At the primary 32768-token selection target, validation-selected mixtures of natural configurations improve pooled accuracy by 1.76 points at 16.86% lower cost than the jointly tuned suffix, meeting the registered accuracy/cost criterion. The other primary comparison, against a fixed-allowance baseline, reduces cost by 19.01% without establishing accuracy superiority; single-configuration selection shows different cost trade-offs. Six historical grids quantify allowance sensitivity, and a plurality-vote account predicts aggregate accuracy with 2.2-point mean absolute error on held-out samples of the same questions. Natural-sampling controls separate prompt-associated discordance from ordinary variation. These results support fairer test-time strategy comparisons that jointly select both allocation controls and report measured costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.