acceptodds
Under review as a conference paper at ICLR 2027

Design Principles for Robust Reinforcement Learning Systems

Abstract

Deep reinforcement learning (RL) algorithms must concurrently coordinate exploration, policy improvement, and neural network optimization. While prior work has seen great success in optimizing these components in isolation, this paper argues that the asynchronous interactions between them are the primary driver of sensitivity and instability in modern deep RL systems. We demonstrate that well-known vulnerabilities, such as batch size limits in DQN, clipping bounds in PPO, and entropy tuning in continuous control, are often symptoms of hyperparameter collusion, where a sub-optimal choice in one component is required to mask inadequacies in another. Drawing an analogy to concurrent software systems, we propose a set of algorithm-agnostic design patterns aimed at ensuring the correctness of each component. We empirically show that these principles significantly improve the robustness and peak performance of widely-used deep RL agents, including PPO and SAC. We hope this work facilitates greater scientific rigor by simplifying standard baselines and streamlining the evaluation of new algorithms

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.