acceptodds
Under review as a conference paper at ICLR 2027

BanditSched: Adaptive Data Scheduling with Contextual Bandits for Efficient Reasoning Reinforcement Learning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting reasoning capabilities in large language models. As the model's competence evolves during training, so does the value of each training problem, and data selection must adapt to this non-stationary process. Existing methods rely on offline profiling, coarse-grained curricula, or filtering after rollouts have already been generated. We observe that the reward-based learning signal of a problem is governed by its current pass rate: groups with intermediate success rates yield nonzero mean absolute advantages, while groups with uniformly correct or incorrect responses yield zero advantages. Motivated by this observation, we formulate RLVR data selection as a *non-stationary contextual bandit* problem, where each training problem is an arm and its context encodes competence statistics, difficulty, and reasoning-pattern information. We propose BanditSched, a Thompson Sampling-based framework that (1) predicts prompt utility before allocating rollouts, (2) periodically probes previously attempted low-success problems to refresh estimates near the model's competence frontier without policy updates, and (3) combines advantage magnitude, competence-edge membership, and question-space diversity into a composite reward. A sliding-window reward model tracks the changing utilities, with its window length guided by an estimation–drift trade-off. Experiments on mathematical reasoning, code generation, and logical puzzles with Qwen2.5 and Qwen3 models show that BanditSched improves rollout efficiency while matching or exceeding the final accuracy of strong baselines, with the largest gains on hard and out-of-distribution tasks. Specifically, BanditSched reaches 95% of Full-GRPO's final accuracy using 30–35% of its online rollouts and completes full training runs with an approximately wall-clock speedup. Controlled comparisons further show that contextual scheduling improves efficiency beyond pass-rate band filtering alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.