RLHF without Reward Clipping: Near-Optimal Distortion at Every KL Budget
Abstract
Reinforcement learning from human feedback (RLHF) fits a single reward model to comparisons from users with heterogeneous preferences and then optimizes a policy under a Kullback–Leibler (KL) constraint. Distortion measures the ratio between the optimal average user utility and that attained by the learned policy. Gölz et al. (2025) show that RLHF can suffer distortion exponential in the Bradley–Terry temperature β, and Oko et al. (2026) trace this growth to distribution mismatch between the preference data and the KL reference policy. Without such mismatch, they prove distortion linear in β, which is the optimal order. However, their guarantee at finite KL budgets relies on clipping the fitted reward, and the best known bound for the unclipped reward is quadratic in β. We show that unclipped RLHF attains this optimal order. Without distribution mismatch and with the reward fitted to population comparisons, unclipped RLHF has worst-case distortion at most C₀κ(β) at every KL budget, where C₀ < 2.526 and κ(β) = β/[2 tanh(β/2)] is the known algorithm-independent lower bound. The main difficulty is that a KL-constrained policy depends on reward gaps as well as reward rankings. We address it by comparing reward maximization with win-probability maximization inside the same KL ball. In population simulations on MovieLens and KuaiRec, the unclipped reward has mean relative utility loss below 0.15% over each dataset's evaluation grid and lower mean loss than the clipping rule of Oko et al. (2026). These results indicate that clipping is also unnecessary on the evaluated user profiles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.