What Governs Learning Dynamics in Massively Parallel Off-Policy RL?
Abstract
GPU-accelerated simulators enable data collection from thousands of environments in parallel for off-policy RL. This parallelism increases data throughput, but also changes the dynamics of the replay buffer distribution during the training. As the replay buffer has finite capacity, increasing data throughput causes its samples to be replaced more rapidly. In this study, we analyze the learning dynamics of reinforcement learning with massively parallel environments. Specifically, we treat the replay buffer as a temporal mixture of data distributions under changing policies and distinguish the true distribution shape from the sample density. We also identify two sources of gradient variance, one induced by replay buffer construction from finite rollout samples and the other by mini-batch sampling from the replay buffer. Our empirical results highlight the importance of increasing the replay buffer size and batch size as the number of parallel environments grows, and reveal regimes in which additional parallelism yields little benefit. These findings provide practical guidance for choosing replay buffer and mini-batch sampling hyperparameters in massively parallel off-policy RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.