It Has the Answer and Keeps Going: How Fine-Tuning Breaks the Decision to Stop
Abstract
Fine-tuning on long reasoning traces is how most open reasoning models are built, and its students often fail to finish: ours runs to the budget on half of its AIME 2025 generations. The cause is a learned stop decision, which fine-tuning can break while leaving reasoning intact. Serve a Qwen3-8B student without its eleven-token, task-free system turn and its computed accuracy moves points while its delivered accuracy falls : it boxes the correct answer early and keeps going. One principle explains every failure we find: a distilled student stops only under conditions that accompanied every stop it was shown. Three causes follow, each confirmed by controlled intervention. (i) An end-of-turn token dropped from supervision is silenced ( of generations) while a standard harness scores both checkpoints alike. (ii) A stop tied to the training prompt fails at the two closing tokens; editing their probabilities restores termination, and training on two formats prevents it. (iii) Every demonstration stops on a converged answer, so a failing student, whose candidate answers never settle, stops after only of its wrong answers (base: ), a pattern of released long-CoT models share. Trained on stops alone, the student collapses; mixed with original data, a moderate dose halves wrong-and-capped generations at no significant accuracy cost on one training seed, and stronger doses abandon solvable problems. Terminator, one guard per cause, terminates non-termination at inference: on every student we trained and benchmark we tested, capping falls to zero and delivered accuracy rises ( points under the training format, without the system turn; also pooled over released checkpoints) while computed accuracy never falls; budget forcing, which also forces a commitment, does as well, as the principle predicts. Reported termination is partly a property of the harness: four standard defaults span points for one checkpoint.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.