BanditSched: Adaptive Data Scheduling with Contextual Bandits for Efficient Reasoning Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting reasoning capabilities in large language models. As the model's competence evolves during training, so does the value of each training problem, and data selection must adapt to this non-stationary process. Existing methods rely on offline profiling, coarse-grained curricula, or filtering after rollouts have already been generated. We observe that the reward-based learning signal of a problem is governed by its current pass rate: groups with intermediate success rates yield nonzero mean absolute advantages, while groups with uniformly correct or incorrect responses yield zero advantages. Motivated by this observation, we formulate RLVR data selection as a *non-stationary contextual bandit* problem, where each training problem is an arm and its context encodes competence statistics, difficulty, and reasoning-pattern information. We propose BanditSched, a Thompson Sampling-based framework that (1) predicts prompt utility before allocating rollouts, (2) periodically probes previously attempted low-success problems to refresh estimates near the model's competence frontier without policy updates, and (3) combines advantage magnitude, competence-edge membership, and question-space diversity into a composite reward. A sliding-window reward model tracks the changing utilities, with its window length guided by an estimation–drift trade-off. Experiments on mathematical reasoning, code generation, and logical puzzles with Qwen2.5 and Qwen3 models show that BanditSched improves rollout efficiency while matching or exceeding the final accuracy of strong baselines, with the largest gains on hard and out-of-distribution tasks. Specifically, BanditSched reaches 95% of Full-GRPO's final accuracy using 30–35% of its online rollouts and completes full training runs with an approximately wall-clock speedup. Controlled comparisons further show that contextual scheduling improves efficiency beyond pass-rate band filtering alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.