Termination Calibration Recovers Answers That Reasoning Models Compute but Never Express
Abstract
Large reasoning models trained with reinforcement learning from verifiable rewards (RLVR) achieve state-of-the-art performance, but they are evaluated under generous token limits while deployment imposes hard budgets. We show that under such budgets, these models fail primarily through ***termination failure*** rather than capability failure: nearly half of a distilled model's truncated generations already contain the correct answer—computed, but never expressed. Existing efficiency methods do not resolve this. Length-penalized rewards teach models to be shorter, yet their stopping points remain coupled to whatever budget is available; and we find the standard correctness-only RLVR recipe can actively worsen termination, lengthening typical traces and dropping tight-budget accuracy below that of the untrained base model. We introduce ***termination calibration*** through **R3**, a single trajectory-level reward granting full credit only when a correct answer terminates within a threshold that tracks the policy's own recent correct-solution lengths. Our design (i) requires no auxiliary judge, token-level labels, or curriculum, and (ii) includes a matched correctness-only control that isolates the causal mechanism. Across budgets and datasets, R3 matches or exceeds specialized efficiency methods at roughly half their tokens and twice the speed; its learned stopping rule is budget-invariant yet extends adaptively to problems harder than any seen in training, and the effect replicates across three backbones. Efficient reasoning, we argue, is less about learning to be brief than learning when to stop.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.