Learning What to Learn: Langevin-based Curriculum Optimization for RLVR
Abstract
Reinforcement learning with verifiable rewards has substantially improved the reasoning capabilities of large language models, but its efficiency is often limited by uniform prompt sampling that ignores differences in training utility under the evolving policy. Curriculum learning offers a natural way to adapt training data to the evolving policy, yet existing methods either rely on predefined difficulty schedules or estimate utility at the individual-prompt level, where sparse rollout feedback can lead to unreliable estimates and self-reinforcing sampling bias. To address these challenges jointly, we propose Langevin-based Adaptive Curriculum Optimization (LACO). LACO groups prompts into buckets based on their initial difficulty and pools online rollout feedback within each bucket, enabling more reliable estimation of their current training utility. It then uses Langevin Monte Carlo to adapt the sampling weights across buckets: utility-driven updates favor currently informative buckets, while stochastic perturbations preserve exploration of under-sampled buckets whose utility may change as the policy learns. This enables LACO to continuously track the policy's moving capability frontier while avoiding premature sampling concentration. Extensive experiments across multiple reasoning benchmarks demonstrate consistent improvements over uniform sampling and existing curriculum methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.