Generative Bellman Flow: Mitigating the CVaR Gradient Bottleneck in Risk-Sensitive Reinforcement Learning
Abstract
Distributional reinforcement learning provides the expressivity to model full return distributions, yet this has not consistently translated into reliable tail-risk policy optimization in continuous control. We identify a CVaR actor-gradient bottleneck: implicit-quantile critics can represent the lower tail, but actor gradients derived from re-sampled tail queries can become noisy when catastrophic and non-catastrophic outcome modes induce well-separated action gradients. To address this, we introduce GBF-SAC, a Soft Actor-Critic instantiation of Generative Bellman Flow. Its conditional flow-matching critic models mean-policy continuation returns and generates a shared sample set for each state-action pair, from which GBF-SAC amortizes the mean and lower-tail statistics into differentiable Q-heads. These heads provide the mean and risk-sensitive actors with smoother scalar objectives rather than repeatedly re-sampled tail estimates. On a suite of risk-sensitive MuJoCo continuous-control tasks with a velocity-dependent crash wrapper, GBF-SAC improves the reward-crash trade-off on the main MuJoCo suite while maintaining competitive mean return. An IQN+Q-Head ablation further isolates the mechanism, indicating that the reward-safety gains arise from the interaction between stable tail representations and the Q-head bridge. Because the critic retains a full learned return distribution, the same trained model also supports diagnostic post-hoc risk scoring at deployment without retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.