The Model is Its Own Curriculum: Internal Difficulty Assessment for Self-Adaptive RL in Reasoning
Abstract
Reinforcement learning can substantially improve language model reasoning, but sparse outcome rewards provide limited guidance for exploring complex reasoning trajectories. Curriculum learning offers a natural way to structure this exploration by adapting difficulty to the evolving policy. A key question, however, is at what granularity such learner-relative difficulty should be defined. While empirical pass rates provide a dynamic estimate of problem difficulty, the current policy's likelihood can further distinguish among correct reasoning trajectories for the same problem. We propose Self-Curriculum, which uses the negative length-normalized log-likelihood of correct trajectories as a learner-relative measure of trajectory difficulty. Self-Curriculum orders verified trajectories from easy to hard, periodically updates their difficulty as the policy evolves, and incorporates this signal into reward shaping. Across multiple mathematical reasoning benchmarks and model families, Self-Curriculum consistently outperforms standard GRPO. On Qwen3-8B, a curriculum based on dynamic pass rates already improves over GRPO, while Self-Curriculum yields further gains and also outperforms the evaluated external difficulty baselines at both the problem and trajectory levels. These results support policy likelihood as a fine-grained, learner-relative signal for adaptive curriculum learning under outcome supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.