acceptodds
Under review as a conference paper at ICLR 2027

The Reward-Aware Laplacian Representation in Zero-Shot Reinforcement Learning

Abstract

Representations based on the reward-agnostic successor measure have been shown to be effective in a variety of applications, and, in particular, have led to many advances in zero-shot RL. In this paper, we consider the setting in which, across different tasks, there are shared reward structures that reflect inherent environmental properties, e.g., the negative rewards of wear and tear in robotics. We demonstrate that in this setting, standard zero-shot RL methods can fail to represent downstream reward functions well, as large approximation errors for the shared reward structures can dominate task-specific reward information during inference. We address this problem by learning a reward-aware representation that reduces the burden of representing shared reward structures during inference, enabling better reconstruction of task-specific rewards. We do so by proposing ReAL, a reward-aware generalization of the Laplacian representation. Empirically, we show that compared to reward-agnostic representations, the ReAL representation enables better downstream reward fitting and stronger zero-shot performance in zero-shot goal-reaching tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.