acceptodds
Under review as a conference paper at ICLR 2027

Reasoning, Answer Delivery, and Budgeted Accuracy

Abstract

A late answer cue can raise a reasoning model's score, but its interpretation depends on the comparison. We compare direct continuation D with a four-token cue L on the same prompt and sampled prefix, under one output cap that includes forced tokens. Correct delivery requires a natural end of sequence (EOS), a complete final boxed answer and acceptance by the fixed scorer, which allows numeric tolerance on MATH. On 200 unused MATH questions, preregistered L raises delivery by 7.33 percentage points, 95% interval [5.00, 10.00], for Qwen3.5-9B at 2048 tokens with 45 rescues and one harm, by 6.83 for Ministral-3-8B and by 4.50 for Qwen3.5-9B at 4096 tokens. Registered GPQA-Diamond transfers give +20.20 [16.50, 24.07] with Qwen and +4.55 [2.69, 6.57] with Ministral without harms, and L generates 1–4% fewer tokens. Only runs active at the cue can change, and across nine settings 23–49% of them convert net of harms. L exceeds a budget-charged variant of the closest protocol's conclusion message by 3.50 points [1.50, 5.67], largely through greater answer room, and a numeric-budget system prompt by 9.50 [6.17, 13.00]. The comparison changes reserve selection: the earliest cue looks best against the stopped stream in three MATH settings, but against continuation to the same cap a 128-token reserve beats it by 15–21 of 600 runs at 2048 tokens. Prospectively, at a 1024-token cap, a 128-token reserve beats a 512-token one by 14.33 [10.67, 18.33] and 10.33 [6.50, 14.33] points for the two models, while matched no-cue continuations deliver within 0.67 points of D. Continuing D also beats restarting thinking at 1024 tokens by 8.25 points at the same cap.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.