Termination Misalignment in Large Reasoning Models
Abstract
We characterize termination misalignment in Large Reasoning Models (LRMs) as the gap between recognizing input unresolvability and acting on that recognition by regulating reasoning to terminate or abstain accordingly. In approximately 73.94% of reasoning trajectories from an evaluation of 23 models over 6,210 inference runs, we observe two distinct behavioral manifestations of the phenomenon: the failure to either terminate reasoning (within 16,384 output tokens), or abstain from hallucinating after terminating reasoning. Crucially, at least 32.77% of these trajectories contained explicit statements recognizing input unresolvability, lack of progress, or drafts of conclusions semantically appropriate for actuating termination. Despite the generation of such statements, models continued to reason for thousands of tokens. We introduce Vocabulary-Based Adversarial Fuzzing (VB-AF), a gray-box fuzzing framework for language models, and use it to generate adversarial, unresolvable inputs to systematically isolate and reproduce this behavior. Context priming models with epistemic principles substantially increases early-termination and abstention behavior, without generally degrading performance on standard tasks or consistently altering the step-level log-odds of continuing to reason. Analysis of primed trajectories suggests that termination is associated with semantically appropriate contextual states rather than a gradual decline of the measured continuation preference. Our findings are consistent with termination misalignment being a reasoning regulation problem, where models can semantically generate termination-relevant statements under persistent unresolvability but still continue deliberation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.