What Do Early Answers Reveal? Readout Protocols and Post-Answer Revision
Abstract
Does a correct continuation from an early reasoning prefix mean the model was ready to stop? We compare free continuation (no extra request) with continuation prompted to give a final answer, scoring both first and last answers. Across Qwen3 and GPT-OSS on MATH-500 and GPQA-Diamond, recovery depends on the request and which answer is scored. On 100 MATH problems, Qwen3-8B with thinking enabled at a 10% prefix achieves 69.3% free last-answer accuracy; requested outputs rise from 31.8% at the first answer to 73.8% at the last. Among improvements, the median delay from the first answer to the first later correct box is 1,258 tokens. Allowing 16,384 reference tokens on 50 reused questions preserves a 52-point first-to-last gain, but no clear last-answer advantage for either procedure. Closing the reasoning channel reduces later gains in several settings, but not clearly for Qwen at 50%. At 50%, equal continuation budgets retain a Qwen prefix benefit with thinking, but not a clear benefit without it. An online case study saves tokens in the main generation while increasing time and total generation. Recovery, prefix benefit, and stopping efficiency therefore require separate evidence: score first and last answers, compare question-only recovery, and count the full cost before interpreting recovery as readiness to stop.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.