acceptodds
Under review as a conference paper at ICLR 2027

Quantile-Constrained Reinforcement Learning via Distributional Credit Assignment

Abstract

In many sequential decision-making applications, such as healthcare and inven- tory control, the goal is not only to achieve high average returns but also to avoid poor lower-tail outcomes. We study expected-return maximization under a lower- quantile constraint and focus on the key challenge of assigning trajectory-level risk to individual decisions. Using an equivalent violation-probability formula- tion, we propose distributional credit assignment with two key components. (1) At each step, we use accumulated rewards to determine how much future return is needed to meet the constraint. We then define a transition-based risk advantage from the change in violation probability between consecutive augmented states. (2) We estimate these violation probabilities from conditional future-return dis- tributions learned through distributional Bellman updates. This uses intermedi- ate reward magnitudes to estimate risk, rather than relying solely on terminal binary outcomes. We combine the resulting risk advantages with reward advan- tages in proximal policy optimization (PPO). Theoretically, we derive a density- free risk-gradient representation, provide error bounds for risk-gradient estima- tion, and establish convergence under a three-time-scale stochastic-approximation formulation. Experiments in synthetic mean-risk trade-off and SimpleButton environments show that our method provides favorable performance in balanc- ing expected return and lower-tail constraints compared with existing quantile- constrained baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.