acceptodds
Under review as a conference paper at ICLR 2027

A Relative Perspective on Distributional Reinforcement Learning

Abstract

How much better is one action than another? We introduce Relative C51 (R-C51), a distributional reinforcement learning method that directly learns distributions of return differences between state–action pairs. A central obstacle is episode termination. When only one of two sampled transitions terminates, subtracting their Bellman targets requires absolute continuation information unavailable from relative comparisons alone. We address this with Bellman uniformization, a transformation that introduces self-loops and rescales rewards to get a common discount factor while preserving expected action values. For a fixed policy, the resulting unprojected operator is contractive, and its unique fixed point recovers the original action-value differences in expectation, although the full distributions can differ from episodic return differences. R-C51 implements this with a categorical relative critic, a learned continuation head, and an explicit mixture over successor comparisons. On the full 57-game Atari benchmark, R-C51 improves on C51 by approximately % in median score, % in interquartile mean, and % in optimality gap. These results support relative distributional prediction as a promising approach to value-based reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.