From Completed Length to Prefix Decisions: A Paired Budget Audit of Reasoning Models
Abstract
A completed response's length can predict correctness while providing little guidance about when to stop generation. We examine this distinction in a paired prefix-replay study of Qwen3-4B-Thinking-2507 and DeepSeek-R1-Distill-Qwen-7B: four sampled streams for each of 300 GSM8K test questions, totaling 2,400 trajectories. An agreement rule stops once two visible numerical answers match; its threshold is chosen on a separate training-split development set. Under a 4,096-token ceiling, agreement scores 58.0% and 59.3%, exactly matching round-robin's question-level correctness. It consumes 131 and 235 fewer tokens per question, but random allocation scores 67.7% and 67.0%. A post hoc control that finishes streams in fixed order reaches 88.3% on both checkpoints. In contrast, retrospective within-question comparisons show longer incorrect than correct responses on the subsets containing both. The results separate that completed-length association from the effect of a prefix-only decision: how the budget reaches a final answer matters more here than stopping on agreement. We report paired uncertainty, all controls, and separate costs for replay and corpus generation. The measured costs count tokens; online serving latency is not evaluated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.