ThinkMid: Thinking and Correcting in the Middle of Response Generation
Abstract
Large Reasoning Models (LRMs) scale at test time by spending more computation on a problem, yet existing methods largely decide how much to think before answering, rather than when to think again while answering. We introduce , a training-free decoding framework that dynamically allocates additional reasoning during response generation. ThinkMid preserves the model’s original reasoning trace, generates the response incrementally, and monitors sentence-level uncertainty using token entropy. When a sentence is sufficiently uncertain, generation pauses and the same model performs an auxiliary self-review in the existing conversational context. If an error is identified, a further reasoning turn produces a local correction; the flagged segment is replaced, and generation resumes from the revised prefix. ThinkMid requires no external verifier or additional training and is complementary to existing test-time scaling strategies. Across mathematical, scientific, and expert-level reasoning benchmarks using two LRMs with different architectures and reasoning formats, ThinkMid improves over the single-solve baseline and compute-matched test-time scaling methods while invoking additional reasoning for only a small fraction of generated sentences. These results indicate that where additional computation is inserted during generation can be as important as how much computation is used.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.