acceptodds
Under review as a conference paper at ICLR 2027

Target Value as Potential: Reward Shaping in Deep Reinforcement Learning by Transferring Prior Knowledge

Abstract

We propose Target Value as Potential (TAP), a simple reward shaping method that accelerates learning in deep reinforcement learning (RL) by transferring prior knowledge. Although reward shaping provides a natural mechanism for leveraging prior knowledge, theoretical analysis and empirical results show that its effectiveness depends on how this knowledge is represented and that priors constructed using existing methods may not reliably improve performance. To address this challenge, TAP introduces a new approach to constructing priors by using only the critic’s target value as the shaping potential within the classical Potential-Based Reward Shaping (PBRS) framework. Consequently, TAP requires neither additional representational structures nor extra hyperparameters. Our analysis and experiments show that TAP accelerates convergence and improves cumulative returns over standard DRL baselines, including TD3 and D4PG. Across tasks in the DeepMind Control Suite, TAP consistently outperforms existing reward shaping methods, such as Dynamic Potential-Based Reward Shaping (DPBRS), and more general methods for incorporating prior knowledge into RL, such as Heuristic-Guided Reinforcement Learning (HuRL). Furthermore, initial results integrating TAP with MR.Q, a representation-learning method for pixel-based observations, also results in improvements across two visual control environments. These results suggest that TAP remains effective even when the shaping potential is derived from high-dimensional learned representations rather than true environment states.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.