Unleashing Flow Policies with Distributional Critics
Abstract
Generative flow models have recently emerged as a powerful paradigm for policy learning in offline reinforcement learning, capable of capturing complex, multimodal action distributions inherent in diverse datasets. However, the full potential of these expressive actors is often bottlenecked by their critics, which typically learn a single scalar estimate of the expected return. We argue that directly regressing this mean value is suboptimal, as it often fails to capture the underlying stochasticity of returns, leading to high-variance and inaccurate value estimates. To address this, we introduce the Distributional Flow Critic (DFC), a novel architecture that utilizes flow matching to learn the complete state-action return distribution. Crucially, we demonstrate that explicitly modeling this distribution acts as a robust representation regularizer: by capturing the comprehensive outcome distribution rather than compressing complex dynamics into a single scalar, DFC yields a much higher-fidelity value signal. We extensively validate our approach across diverse tasks from the D4RL and OGBench benchmarks, spanning locomotion and manipulation. Results show that DFC achieves the highest score on out of domain groups, with an average domain-wise score ratio of relative to the strongest baseline in each group. DFC further excels in offline-to-online adaptation, achieving the best or tied-best fine-tuned performance on out of domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.