DRL Plus: A Replay-Free and Non-Bellman Policy-Learning Framework
Abstract
Reinforcement Learning (RL) has become a widely used framework for sequential decision-making, particularly in robotics, autonomous systems, and complex control tasks. Recent advances in Deep Reinforcement Learning (DRL) have enabled learning in high-dimensional environments; however, many conventional approaches rely on Bellman-based value estimation, carefully designed reward functions, and, particularly in off-policy methods, experience replay buffers. These dependencies can contribute to challenges such as high sample complexity, sensitivity to reward design, training instability, memory overhead, and safety concerns associated with trial-and-error learning in real-world applications. In this paper, we propose DRL Plus, a framework formulated within the Markov Decision Process (MDP) setting that eliminates Bellman-based value function estimation, handcrafted reward functions, and experience replay. Safety and performance are critical in robotics and autonomous systems; therefore, autonomous driving is selected as the target domain due to its complexity and stringent safety requirements. Through formal analysis and a set of carefully designed challenging driving scenarios, we demonstrate that DRL Plus achieves improved performance and safety compared with state-of-the-art deep reinforcement learning and safe reinforcement learning approaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.