Undiscounted Potential Based Reward Shaping
Abstract
Reinforcement learning usually trains with a discount factor for stability, even when the task's objective is the undiscounted return. In this paper, we compare two forms of a frequently used reward shaping method, which adds a shaping term to immediate rewards to speed up learning. Potential-based reward shaping (PBRS) provides a framework which guarantees that the shaping term does not change the discounted optimization objective. However, when training uses , the undiscounted objective does not necessarily coincide with the discounted objective. This means that PBRS cannot shape rewards towards an undiscounted optimum when the discounted objective differs from the undiscounted one. To provide stable training in these cases, we examine undiscounted potential-based reward shaping (PBRS), which drops the discount factor from the shaping term but still uses discounted targets when training reinforcement learning agents. Though PBRS drops the guarantee to be objective-preserving, we show that when the potential function corresponds to the optimal state value function, the PBRS-shaped rewards correspond to the action advantage function of the undiscounted objective and reproduce an optimal policy. As in practice we do not have optimal values, we further present an error bound based on the estimation bias of the state value function. In our experiments, PBRS beats PBRS and the correction term shaping with the same learned potential in several settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.