Probing Reward Beliefs in Reinforcement Learning
Abstract
Driven by a decade of rigorous research, deep neural policies can now navigate complex MDPs, learn how to reason and strategize in high dimensional environments. While reinforcement learning is now a cornerstone of varied domains, spanning from language model mathematical and scientific reasoning to sensitive settings such as medical and finance, a prominent line of research seeks to infer reward functions by observing how an optimal policy behaves. This approach is predicated on the claims that reconstructing the underlying objective leads to policies more closely aligned with intended outcomes than those defined by hand. In this paper, we analyze the underpinnings of learning with reward beliefs in high-dimensional state representation MDPs. We demonstrate that this premise has fundamental shortcomings and standard deep reinforcement learning yields more resilient and value-aligned policies when compared to learning from the behaviour of other policies in MDPs with complex state representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.