Rollout Aggregation Stabilizes On-Policy Reinforcement Learning
Abstract
Large-scale parallel PPO can struggle on complex, long-horizon tasks with high-dimensional action spaces. Prior work has emphasized exploration bottlenecks and diminishing returns from increasing environment parallelism. We instead present evidence that learning instability offers a compelling alternative explanation for PPO's poor performance. To improve stability, we introduce rollout aggregation (RAgger), a minimal modification that retains the previous rollout and trains on it alongside newly collected samples. We show that overlapping rollout batches regularize the policy's optimization trajectory by preserving pressure to retain recently learned behavior while incorporating new experience. Across several challenging manipulation tasks, RAgger prevents the collapse observed with standard PPO and outperforms strong baselines, including SAPG and CPO, with several-fold improvements in sample efficiency. Our findings motivate studying recent-data reuse as a mechanism for stabilizing policy learning and achieving strong, sample-efficient performance at scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.