acceptodds
Under review as a conference paper at ICLR 2027

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

Abstract

Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how they are weighted, and which domains contribute to subsequent batches. We introduce DataFlex-RL, an evaluation platform for comparing these policies under a common GRPO recipe. In our primary study, we evaluate 13 configurations with 12 matched seeds on Qwen2.5-7B-base across 12 math, logic, and science benchmarks. Uniform GRPO improves domain-balanced mean accuracy (Overall) by 7.76 percentage points over the untrained checkpoint. For none of the eight selection or reweighting methods does the paired 95% confidence interval for the difference from uniform sampling exclude zero. Likewise, none of the three adaptive mixtures significantly outperforms a fixed equal-weight mixture at this precision. A corrected 12-seed extension on Llama-3.1-8B-base places the additional methods on the same score scale as the original controls, but their observed means reveal no consistent winner. To assess evaluation sensitivity, we recompute scores for nine Qwen2.5-7B-Instruct runs using both a math-heavy six-benchmark summary—five math benchmarks plus GPQA-Diamond, excluding logic—and the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated (), whereas rankings based on summaries retaining all 12 benchmarks are largely consistent. Overall, across the controlled settings studied, data policies measurably alter the training process but do not provide a reproducible advantage over uniform training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.