Bellman residual minimization for sample-efficient deep reinforcement learning
Abstract
Recent work on improving the sample efficiency of deep reinforcement learning has largely focused on network architecture design, normalization schemes, and optimizer choices. The loss function itself, however, has received comparatively little attention, even though it plays a central role in shaping training dynamics. We show that enforcing an additional Bellman consistency objective on top of the semi-gradient loss stabilizes training at high replay ratios. Our method demonstrates competitive performance with concurrent methods on the hardest DeepMind Control Suite tasks like humanoid and dog, without relying on network resets or additional normalization techniques.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.