Can Language Models Learn to Reject Their Own Bad Reasoning Steps?
Abstract
Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. Can a language model instead learn to reject its own bad reasoning steps? This requires distinguishing candidate steps sampled from the same problem and reasoning prefix. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostic measurements show that most adjacent changes in recoverability are difficult to resolve with practical Monte Carlo budgets, even when cumulative degradation is detectable. Meanwhile, candidates sampled from a shared prefix exhibit a sparse low-recoverability tail, suggesting an opportunity for selective rejection. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator's own backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are labeled as accepted, the crossing step is labeled as rejected, and unresolved steps and the subsequent suffix are excluded from supervision. Training combines pointwise classification with same-prefix pairwise learning, followed by group-relative policy refinement using final-answer correctness. At inference, the gate accepts candidates or resamples from the unchanged prefix, subject to a per-position attempt cap and an episode-level rejection budget, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4–10.1 percentage points using 1.21–1.40 as many generated tokens, achieving the highest macro-average accuracy among the evaluated step-level methods for each generator, while full-solution scaling methods require 4.47–8.27 the single-pass token cost for comparable performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.