acceptodds
Under review as a conference paper at ICLR 2027

Safe RLHF Beyond Expectation: Stochastic Dominance for Spectral Risk Control

Abstract

Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety by constraining the expected cost of a policy. However, expected cost captures only average behavior and can overlook rare but severe failures. We propose Risk-sensitive Alignment via Dominance (RAD), a distributional approach to safe RLHF that compares the cost distribution of the learned policy to that of a reference policy rather than only their means. RAD represents policy-induced cost distributions non-parametrically through quantiles and encourages the learned policy to have lower cost quantiles than the reference policy. To optimize this distributional objective for language models, RAD uses a quantile-particle representation that connects two gradient estimators: entropic optimal transport provides gradients of the distributional objective with respect to quantile particles, while score-function quantile gradients propagate these updates to the policy. We further generalize RAD with quantile weighting, allowing different regions of the cost distribution to be emphasized. For admissible spectral weights, the resulting objective has a direct connection to Spectral Risk Measures, providing a principled way to target different notions of risk. Empirically, on BeaverTails, RAD improves harmlessness over SFT, and several variants also improve harmlessness over Safe RLHF while maintaining similar helpfulness. On held-out HarmBench prompts, several risk-sensitive variants also improve harmlessness over Safe RLHF.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.