Learning to Redirect and Stop: Self-Teaching Control for Reasoning Models
Abstract
Large language models (LLMs) have demonstrated strong reasoning capabilities through extended chain-of-thought generation, yet they may take suboptimal directions at critical reasoning transitions or continue reasoning after a correct answer has already emerged. We introduce Self-teaching Decision and Exit for Reasoning (SeDER), a lightweight framework that adaptively controls both the direction and length of reasoning while keeping the underlying language model frozen. At its core, a decision-token head selectively intervenes at uncertain reasoning transitions. Using successful rollouts generated by the base model itself, we construct counterfactual alternatives and evaluate them based on their compatibility with successful continuations and stability under short stochastic rollouts. The resulting utility differences yield weighted pairwise preferences for training the decision-token head to rerank plausible reasoning operators, without requiring a stronger teacher, a separate process reward model, or reinforcement learning. We further incorporate a lightweight stopping head that terminates reasoning once a sufficient solution has been reached. Experiments across multiple reasoning benchmarks with DeepSeek-R1-Distill-Qwen-1.5B and 7B demonstrate consistent accuracy improvements while substantially reducing reasoning length. SeDER achieves the highest accuracy among the evaluated baselines across all benchmarks, while ablation studies validate the effectiveness of its utility construction and selective intervention strategy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.