acceptodds
Under review as a conference paper at ICLR 2027

Flow-time Bellman Matching for Offline Distributional Reinforcement Learning

Abstract

Distributional reinforcement learning models the full return distribution rather than only its expectation, providing a richer representation of future returns for value learning. Flow matching provides a flexible framework for modeling return distributions, but existing approaches typically impose Bellman consistency only on the terminal return distribution, leaving the temporal structure of the transport process unused. We introduce Flow-time Bellman Matching (FBM), which extends Bellman regression from the terminal return distribution to the entire return flow by constructing Bellman targets from the successor flow at each corresponding flow time. We show that the proposed flow-time Bellman operator is consistent with the standard distributional Bellman operator at the endpoint. To enable efficient evaluation, we further distill the return flow into a one-step quantile critic using an exact correspondence between Gaussian source noise and return quantiles. Experiments on OGBench and D4RL show that achieves the best or near-best performance on 25 of 37 state-based tasks, while controlled analyses demonstrate the contribution of flow-time Bellman supervision and quantile distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.