acceptodds
Under review as a conference paper at ICLR 2027

Same Beliefs, Different Policies: Learning Reusable Representations for Early Stopping in LLM Reasoning

Abstract

Reasoning LLMs often continue reasoning after reaching a correct answer, wasting computation and sometimes changing it to an incorrect one. Deciding when to stop is a metareasoning problem: the model must weigh what further reasoning could achieve against its cost, given what it can already answer. We introduce BELT (BELief-based Termination), which learns beliefs about current and future answers from a frozen LLM's hidden states, then fits a linear stopping policy. We show that predicting reasoning outcomes yields reusable features for stopping control. Changing the token price changes the value of these outcomes, but leaves their prediction targets unchanged, so only the policy needs refitting. Under common policy fitting, belief-trained representations yield higher stopping utility than raw or utility-trained features. Averaged over five out-of-distribution math benchmarks, BELT reduces generated tokens by 66% on Qwen3-1.7B and 53% on OLMo-3-7B-Think at matched test-frontier accuracy relative to natural completion. The same math-trained probe transfers to coding without code-side training. In both domains, its accuracy-cost trade-offs are competitive with or better than those of end-to-end controllers. On Qwen3, it uses 6.3x fewer controller parameters than the strongest baseline and  25x less training compute across five token prices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.