acceptodds
Under review as a conference paper at ICLR 2027

ThinkMid: Thinking and Correcting in the Middle of Response Generation

Abstract

Large Reasoning Models (LRMs) scale at test time by spending more computation on a problem, yet existing methods largely decide how much to think before answering, rather than when to think again while answering. We introduce , a training-free decoding framework that dynamically allocates additional reasoning during response generation. ThinkMid preserves the model’s original reasoning trace, generates the response incrementally, and monitors sentence-level uncertainty using token entropy. When a sentence is sufficiently uncertain, generation pauses and the same model performs an auxiliary self-review in the existing conversational context. If an error is identified, a further reasoning turn produces a local correction; the flagged segment is replaced, and generation resumes from the revised prefix. ThinkMid requires no external verifier or additional training and is complementary to existing test-time scaling strategies. Across mathematical, scientific, and expert-level reasoning benchmarks using two LRMs with different architectures and reasoning formats, ThinkMid improves over the single-solve baseline and compute-matched test-time scaling methods while invoking additional reasoning for only a small fraction of generated sentences. These results indicate that where additional computation is inserted during generation can be as important as how much computation is used.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.