SDRB: A Synthetic Benchmark for Distributional Reinforcement Learning
Abstract
Distributional reinforcement learning (DRL) models complete return distributions, yet standard benchmarks primarily evaluate control performance rather than the fidelity of the learned distribution. We introduce the Synthetic Distributional Reinforcement Learning Benchmark (SDRB), a controlled testbed for directly evaluating return-distribution reconstruction against exact or high-precision ground truth. SDRB isolates six statistical regimes: Gaussian, mixture, heavy-tailed, sparse, distribution drift, and zero-shot out-of-distribution evaluation. We demonstrate the benchmark through a 540-run factorial study crossing three quantile representations (QR-DQN, IQN, and FQF) with three learning objectives (Quantile Huber, MMD, and Sinkhorn divergence). The experiments reveal substantial representation–objective interactions, reconstruction errors in tail and point-mass structure that are not fully captured by global metrics, failure to recover under persistent stale replay after distribution drift, and mean-preserving OOD shifts that substantially increase distributional error while leaving expected-value error essentially unchanged. These results demonstrate how controlled reconstruction evaluation can expose behavior obscured by aggregate control or expected-value metrics. SDRB complements complex control benchmarks by treating return-distribution reconstruction as a distinct evaluation axis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.