acceptodds
Under review as a conference paper at ICLR 2027

From Completed Length to Prefix Decisions: A Paired Budget Audit of Reasoning Models

Abstract

A completed response's length can predict correctness while providing little guidance about when to stop generation. We examine this distinction in a paired prefix-replay study of Qwen3-4B-Thinking-2507 and DeepSeek-R1-Distill-Qwen-7B: four sampled streams for each of 300 GSM8K test questions, totaling 2,400 trajectories. An agreement rule stops once two visible numerical answers match; its threshold is chosen on a separate training-split development set. Under a 4,096-token ceiling, agreement scores 58.0% and 59.3%, exactly matching round-robin's question-level correctness. It consumes 131 and 235 fewer tokens per question, but random allocation scores 67.7% and 67.0%. A post hoc control that finishes streams in fixed order reaches 88.3% on both checkpoints. In contrast, retrospective within-question comparisons show longer incorrect than correct responses on the subsets containing both. The results separate that completed-length association from the effect of a prefix-only decision: how the budget reaches a final answer matters more here than stopping on agreement. We report paired uncertainty, all controls, and separate costs for replay and corpus generation. The measured costs count tokens; online serving latency is not evaluated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.