Starting Near the Goal: Curriculum Reinforcement Learning for Training Initially Weak LLM Policies
Abstract
Reinforcement learning with verifiable rewards is difficult for initially weak language-model policies. An initially weak policy may rarely discover successful trajectories, and terminal rewards provide little direct guidance for intermediate reasoning decisions. To address these challenges, we propose Curriculum Reinforcement Learning (CRL), a training framework that begins exploration near the goal state and uses outcome feedback to adjust curriculum. Starting with shorter continuations reduces the difficulty of obtaining successful outcomes early in training, helping to alleviate the RL cold-start problem. Adaptive curriculum strength and infix reinforcement then support policy improvement by adjusting external guidance and reinforcing reasoning patterns from the policy’s own successful experience. Together, these mechanisms make curriculum responsive to the needs of RL and facilitate the transition to autonomous reasoning as reference guidance is gradually withdrawn. Extensive evaluations across multiple mathematical reasoning benchmarks show that CRL improves both reasoning performance and training efficiency. The code is available at: https://anonymous.4open.science/r/LLM-CRL
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.