Constructing Pareto Frontiers of Optimal Reasoning Termination Methods for Overthinking in Reasoning Models
Abstract
Reasoning models can benefit from test-time scaling, but continued reasoning may become ineffective or harmful, a phenomenon termed overthinking. Evaluations against full thinking compare accuracy and token usage but lack references for overthinking severity and strategies’ theoretical mitigation potential. To characterizes overthinking at the model–task level, we first introduce Step-wise Truncation Probing (STP), repeatedly generating answers from multiple retained prefixes along each fixed sampled reasoning trajectory. Using each trajectory’s earliest measured position attaining maximum empirical accuracy as a reference, we define Overthinking Ratio and Accuracy Degradation as the fraction of thinking tokens generated afterward and the accuracy decrease from this reference to full thinking, respectively. Secondly, we then formulate dynamic routing, fixed budgets, and dynamic early exit within a unified framework to quantify the theoretical mitigation potential of different termination strategies. Based on this unified framework, we use the complete STP measurements to construct a token–accuracy Pareto frontier for each strategy under a common normalized thinking-token budget. We further develop OTBench, a reusable benchmark built from the STP process, which produced 70.5B reasoning tokens and the corresponding probing measurements across six models and four benchmark groups.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.