A Relative Perspective on Distributional Reinforcement Learning
Abstract
How much better is one action than another? We introduce Relative C51 (R-C51), a distributional reinforcement learning method that directly learns distributions of return differences between state–action pairs. A central obstacle is episode termination. When only one of two sampled transitions terminates, subtracting their Bellman targets requires absolute continuation information unavailable from relative comparisons alone. We address this with Bellman uniformization, a transformation that introduces self-loops and rescales rewards to get a common discount factor while preserving expected action values. For a fixed policy, the resulting unprojected operator is contractive, and its unique fixed point recovers the original action-value differences in expectation, although the full distributions can differ from episodic return differences. R-C51 implements this with a categorical relative critic, a learned continuation head, and an explicit mixture over successor comparisons. On the full 57-game Atari benchmark, R-C51 improves on C51 by approximately % in median score, % in interquartile mean, and % in optimality gap. These results support relative distributional prediction as a promising approach to value-based reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.