acceptodds
Under review as a conference paper at ICLR 2027

ProbeCascade: Verification and Value for Early Exit in Long Form Reasoning

Abstract

In deployed long-form reasoning systems, models may continue generating substantial intermediate reasoning after a sufficiently reliable answer has emerged, directly increasing response latency and inference cost. Reasoning early stopping therefore requires recovering as much unnecessary computation as possible while keeping the accuracy loss introduced by early exits within an acceptable budget. Existing methods infer when to stop from local trajectory evidence such as answer stability, confidence, semantic redundancy, or internal representations, but strong local evidence can still coincide with an incorrect candidate and, even when reliability is high, does not by itself determine whether the current checkpoint is a valuable place to exit. We introduce ProbeCascade, a sentence-level early-stopping framework for accuracy-constrained computation recovery that explicitly separates checkpoint discovery from the decision to return the current candidate. ProbeCascade uses local convergence to trigger checkpoint inspection, estimates the reliability of the exposed candidate with a calibrated verifier, and combines this estimate with a Value signal that favors exits from which more computation can be recovered. It stops at the first eligible checkpoint whose score exceeds the selected threshold. To make efficiency gains reflect deployable savings rather than gross truncation, we charge every forced-answer probe actually issued by the stopping policy and report audited net token saving after monitoring overhead. We evaluate ProbeCascade on DeepSeek-R1-Distill-Qwen-7B with MATH-500, DeepSeek-R1-Distill-Llama-8B with MATH-500, and DeepSeek-R1-Distill-Qwen-7B with GSM8K under a common held-out evaluation protocol. On the primary Qwen-7B/MATH-500 setting, the policy selected under a nominal 1.0% accuracy-loss budget achieves 32.06% audited net token saving with 1.2 percentage points of observed accuracy loss, compared with 25.66% saving for Dynasor at the same observed-loss cap, while ProbeCascade also achieves positive audited computation recovery on both additional model–task settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.