acceptodds
Under review as a conference paper at ICLR 2027

Hamilton-Jacobi-Bellman Constrained Flow-based Distributional Q-Learning

Abstract

Flow-based distributional critics learn expressive return distributions from transition-level Bellman supervision, but this supervision constrains the local derivatives of the induced value function only indirectly. We introduce HJB-Constrained Value Flows, which augments this learning process with a stationary Hamilton–Jacobi–Bellman (HJB)-inspired differential regularizer. The noise-averaged initial velocity provides an expected-return estimate, motivated by the population flow-matching regression, without integrating the flow to its terminal time. We use this interface to motivate a local prior from an auxiliary unit-cost control problem and implement it through a single-noise stochastic surrogate. The return-flow architecture, distributional targets, and policy-extraction procedure remain unchanged. Experiments on 25 OGBench manipulation tasks and 12 D4RL Adroit tasks show improvements over the base Value Flows critic, with task-dependent gains. Five offline-to-online evaluations compare the complete training pipelines. Value-landscape and execution analyses illustrate accompanying changes in the critic and policy, while return-spread diagnostics show that the learned distributions retain nontrivial variability. These comparisons support the utility of the complete regularizer for flow-based value learning; they do not isolate its HJB-specific mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.