Curriculum Learning Teaches LLMs Long-Horizon Implicit Reasoning, Up To A Limit
Abstract
LLMs “reason” by externalizing intermediate steps as tokens through chain-of-thought (CoT), yet even frontier reasoning models hit abrupt performance cliffs on long-horizon tasks. CoT commits to one step at a time, and the model's implicit reasoning—reasoning within a single forward pass—sets how far the model can search ahead to decide the correct next move. Models with weak implicit reasoning are more likely to commit to wrong paths and must attempt to recover via backtracking, while models with strong implicit reasoning identify the correct proof strategy directly, shortening the CoT. However, prior work shows LLMs struggle to reason implicitly beyond shallow depths, with no relief from scale. We investigate whether curriculum learning can extend this depth by fine-tuning on gradually increasing vocabulary size and lookahead—the number of steps a model must search ahead to identify the correct next move. At matched compute, curriculum-trained models achieve lookaheads where no-curriculum training fails—Qwen3-0.6B reaches accuracy at lookaheads up to with curriculum, compared to without. However, progress slows with training compute, following a saturating exponential that suggests a maximum intrinsic lookahead to which models can learn to reason implicitly. Curriculum step sizes that are too large stall progression, larger models reach higher fitted ceilings, and pretrained initialization far outpaces training from scratch. Finally, the capability acquired on synthetic data transfers to natural language reasoning, with curriculum-trained models outperforming models trained without curriculum at matched compute on natural language reasoning benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.