Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
Abstract
In the era of deep reinforcement learning, making progress is more complex, as the collected experience must be compressed into a deep model for future exploitation and sampling. Many papers have shown that training a deep learning policy under changing state and action distributions leads to suboptimal performance, or even collapse. This naturally raises the concern that even if the community develops improved exploration algorithms or reward objectives, will those improvements fall on the deaf ears of optimization difficulties? This work proposes a new pracitcal sub-optimality estimator to determine optimization limitations of deep reinforcement learning algorithms. Through experiments acrossenvironments and RL algorithms, it is shown that the difference between the best data generated is better than the policies' learned performance. This large difference indicates that deep RL methods only exploit half of the good experience they generate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.