acceptodds
Under review as a conference paper at ICLR 2027

How Faithful Is a Truncated Reasoning Trace? Measuring and Modeling Biased Feedback in Test-Time Compute Allocation

Abstract

Test-time compute allocation methods must decide how many reasoning tokens to spend on each query. A tempting shortcut is to truncate one long reasoning trace and score each prefix, which yields feedback for every smaller budget without another rollout. Whether such truncated observations faithfully stand in for a native short run has not, to our knowledge, been measured. Comparing them against a same-arm decoding-noise control across budgets from 128 to 2048 output tokens, we find no detectable excess disagreement at budgets of 256 tokens or more, but an excess at 128 tokens (22.5% vs. 11.3% disagreement), the regime where a cost-aware allocator most wants to operate. The excess is bidirectional, so no fixed correction removes it. In an offline replay of an online allocator, reusing truncated feedback helps when tokens are expensive and shows no detectable benefit when they are cheap; removing the 128-token budget from the ladder makes reuse beneficial at the cost weights we tested. Theoretically, we show that pooling biased free observations can be inconsistent, give a regret bound that adapts to which levels’ free observations are trusted, and show that each level’s free observations saturate, so they cannot replace native pulls. Our results come from one model and 75–160 problems, and no learned policy we tested saves tokens at matched top quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.