When to Stop Thinking: Answer-Movement Certificates and Conditional Regret Guarantees
Abstract
Reasoning language models can improve their answers by thinking longer, but additional reasoning consumes resources and can also move an answer in the wrong direction. Efficient inference therefore requires deciding when to stop without access to the correct answer. We develop a framework that connects this decision to changes in the distribution of answers the model would give if stopped now. We derive answer-movement certificates that bound changes in any bounded answer score without correctness labels, yielding conditional regret bounds under cumulative movement budgets and sharper bounds near an optimal stopping policy. The analysis accommodates both beneficial and harmful reasoning and establishes why entropy stabilization alone is insufficient. We complement the theory with a calibrated entropy controller evaluated on 463 questions across four mathematics benchmarks using Qwen3.5-9B and Qwen3.5-27B. Across all eight model–benchmark comparisons, the controller reduces mean thought-plus-answer output tokens by 25–62% relative to horizon-limited native termination. Measured accuracy losses on competition tasks and the consistently lower cost of calibrated fixed budgets expose the limits of the entropy proxy. The framework provides explicit conditions for reliable stopping and a practical methodology for testing whether inexpensive stopping signals deliver useful accuracy–cost tradeoffs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.