acceptodds
Under review as a conference paper at ICLR 2027

Learning to Redirect and Stop: Self-Teaching Control for Reasoning Models

Abstract

Large language models (LLMs) have demonstrated strong reasoning capabilities through extended chain-of-thought generation, yet they may take suboptimal directions at critical reasoning transitions or continue reasoning after a correct answer has already emerged. We introduce Self-teaching Decision and Exit for Reasoning (SeDER), a lightweight framework that adaptively controls both the direction and length of reasoning while keeping the underlying language model frozen. At its core, a decision-token head selectively intervenes at uncertain reasoning transitions. Using successful rollouts generated by the base model itself, we construct counterfactual alternatives and evaluate them based on their compatibility with successful continuations and stability under short stochastic rollouts. The resulting utility differences yield weighted pairwise preferences for training the decision-token head to rerank plausible reasoning operators, without requiring a stronger teacher, a separate process reward model, or reinforcement learning. We further incorporate a lightweight stopping head that terminates reasoning once a sufficient solution has been reached. Experiments across multiple reasoning benchmarks with DeepSeek-R1-Distill-Qwen-1.5B and 7B demonstrate consistent accuracy improvements while substantially reducing reasoning length. SeDER achieves the highest accuracy among the evaluated baselines across all benchmarks, while ablation studies validate the effectiveness of its utility construction and selective intervention strategy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.