SteadyStop: Risk-Aware Early Stopping for Efficient Reasoning
Abstract
Long chains of thought can improve problem solving, but models may continue generating after a useful answer is already available. Stopping earlier can reduce this redundancy, yet an interruption may discard reasoning needed to reach a correct answer. Efficient stopping therefore requires distinguishing readiness to answer from the risk of interrupting an otherwise successful trajectory. To address this problem, we propose SteadyStop, a lightweight stopping framework with separately learned Readiness and Hazard signals. Readiness is trained to predict whether all sampled early answers are correct, while Hazard targets unreliable early answering when full reasoning produces a correct answer. Both signals are learned from reference-checked outputs of the same model using two linear heads with 66 trainable parameters. A Steadiness condition reuses Readiness from the preceding anchor, and the first passing anchor triggers one final answer. The reasoning model remains unchanged, with no auxiliary language model required for annotation or online verification. Across two reasoning models and three mathematical benchmarks, SteadyStop reduces output length by 22.86% with a 0.21 percentage point decrease in accuracy, averaged equally over the six settings. This yields a larger mean token reduction and a smaller mean accuracy decrease than the four evaluated baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.