acceptodds
Under review as a conference paper at ICLR 2027

Learning to Reason More Implicitly Improves Explicit Reasoning on Harder Problems

Abstract

Large language models (LLMs) trained with chain-of-thought (CoT) on simpler problems often fail to generalize to harder problems that require more reasoning steps. Humans, by contrast, can learn to solve harder problems by practicing on simpler problems until they can perform known steps mentally. Inspired by this, we ask whether LLMs can also generalize to harder problems by learning to reason more implicitly, i.e., to perform more steps without generating intermediate tokens, even when they still generate CoTs at test time. In this work, we propose step dropout (SeDo), a simple approach for teaching models to reason more implicitly. SeDo trains models on both complete CoTs and CoTs with randomly omitted steps. This teaches models individual steps while also encouraging them to implicitly infer missing CoT steps. Across five models and four tasks, our results show that SeDo both (1) substantially improves the learning efficiency of implicit multi-hop reasoning, and (2) helps models generalize to harder problems. Surprisingly, SeDo models generalize better even when given complete CoT prefixes, indicating that implicit reasoning also teaches more generalizable skills, beyond robustness to incomplete steps. We further show implicit reasoning helps explicit reasoning in models not trained with SeDo: in looped Transformers, more loops, which allow more computation for implicit reasoning, monotonically improve single-step accuracy within a CoT, even on calculations the models already solve perfectly without CoT.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.