acceptodds
Under review as a conference paper at ICLR 2027

Supervise What You Certify: Stop-Harm Prediction for Early Stopping in Reasoning Models

Abstract

Reasoning models gain accuracy by generating long chains of thought, but incur substantial inference cost and often continue reasoning after reaching a correct answer. Early stopping aims to remove redundant computation while limiting accuracy loss. Existing learned stoppers use supervision based on current-answer correctness, answer agreement, or later answer stability. These targets do not directly capture lost-correct events: stopping yields a wrong answer although full reasoning would answer correctly. The probability of this event upper-bounds the accuracy drop, yet the loss used for certification differs from these supervision targets. We call this event *stop harm* and introduce StopHarm to supervise it directly at each checkpoint, aligning the learning target with the policy loss at the actual stopping point. StopHarm learns a lightweight predictor from paired stopping and full-reasoning outcomes offline, and stops online when the predicted stop harm stays below a certified threshold. Independent request-level certification bounds lost-correct risk, guaranteeing an accuracy drop of at most with confidence under i.i.d. sampling. We prove that even Bayes-optimal proxy predictions can discard stop-harm information that no score-only post-processing can recover. Across three reasoning models and five benchmarks, StopHarm reduces generated tokens by **25.4–38.3%** with average accuracy drops of only **0.27–0.73** points and **1.71–1.77** single-request speedups.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.