acceptodds
Under review as a conference paper at ICLR 2027

Half the Rollouts, No Measurable Cost: A Seeded Audit of Dynamic Sampling and Credit Assignment in DAPO

Abstract

Reinforcement learning for mathematical reasoning is routinely evaluated on AIME, where a single problem is worth over three points, and improvements of one to two points are commonly reported from a single training run. We ask whether such improvements are measurable, and answer with a pre-registered, seeded audit of DAPO on Qwen2.5-Math-7B. Three seeds of the unmodified baseline, read out at 128 samples per problem on AIME 2024 and 2025 pooled, span 1.96 points with a standard deviation of 0.99; across thirteen single-seed runs the rank correlation between the two AIME years is −0.03. Against that unit, we built nine variants: token-level credit-assignment methods that reweight advantages by pivotality, add contrastive step rewards, or importance-weight positions by future influence, and a prompt selector driven by a tracked difficulty estimate; none is individually distinguishable from the baseline under its three-seed prediction interval, and the modified-recipe runs sit on average about 1.5 points below it. A KL-targeted step-size controller, holding each variant at the baseline's late-phase KL, does not recover the largest harm, from high-amplification reweighting. The positive finding concerns the rollout budget: the 75% of prompts that DAPO's dynamic sampling discards are mostly surplus truncated once the training batch is filled; under a trained policy the mixed-group rate is 60–76%, far above the 25% implied by the discard statistic. Halving the generation batch trains the same 64 groups per step at half the rollouts and 24% less wall-clock, with a pooled accuracy difference of −0.18 points over three seeds each, inside the seed standard deviation on both sides.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.