acceptodds
Under review as a conference paper at ICLR 2027

DRL Plus: A Replay-Free and Non-Bellman Policy-Learning Framework

Abstract

Reinforcement Learning (RL) has become a widely used framework for sequential decision-making, particularly in robotics, autonomous systems, and complex control tasks. Recent advances in Deep Reinforcement Learning (DRL) have enabled learning in high-dimensional environments; however, many conventional approaches rely on Bellman-based value estimation, carefully designed reward functions, and, particularly in off-policy methods, experience replay buffers. These dependencies can contribute to challenges such as high sample complexity, sensitivity to reward design, training instability, memory overhead, and safety concerns associated with trial-and-error learning in real-world applications. In this paper, we propose DRL Plus, a framework formulated within the Markov Decision Process (MDP) setting that eliminates Bellman-based value function estimation, handcrafted reward functions, and experience replay. Safety and performance are critical in robotics and autonomous systems; therefore, autonomous driving is selected as the target domain due to its complexity and stringent safety requirements. Through formal analysis and a set of carefully designed challenging driving scenarios, we demonstrate that DRL Plus achieves improved performance and safety compared with state-of-the-art deep reinforcement learning and safe reinforcement learning approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.