acceptodds
Under review as a conference paper at ICLR 2027

From Discovery to Mastery: Reallocating the Replay Budget in Massively Parallel Reinforcement Learning

Abstract

Massively parallel interaction makes rare successful behavior available for learning, yet discovering these behaviors does not ensure that a policy learns to reproduce them. We identify a discovery–mastery gap: policies learn to reliably reproduce discovered high-quality behaviors only after substantial delays or fail to do so within the training budget. We find that increasing the replay share of successful experience can accelerate learning and enable task success without increasing interaction or update budgets, suggesting that insufficient exposure to such experience contributes to this gap. Motivated by these findings, we propose return-guided reallocation of the replay budget (REAP), a generic sampling mechanism that mixes a base sampling distribution with one proportional to positive centered return-to-go. Its return-based scores are computed once from observed rewards when each episode ends and remain fixed as training proceeds. Across 48 control tasks, REAP enables learning on challenging sparse-reward tasks where baselines struggle and substantially improves sample efficiency on the remaining tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.