StableMax-PPO: Best-of- RLHF with Ordinal Reward Distributions
Abstract
Many language-model systems use Best-of- sampling at deployment, while standard reinforcement learning from human feedback (RLHF) optimizes the expected reward of a single response. Existing Best-of- training objectives often reduce each response to a scalar and discard the per-rating probabilities that determine the group maximum. This matters because two responses with the same predicted mean can have different probabilities of receiving a high rating. We introduce StableMax-PPO, which directly optimizes the expected best latent rating among candidates using the predictive distribution of a frozen ordinal reward model trained from human ratings. Its variance is retained as a model-based diagnostic of annotator disagreement, not an optimization bonus. A one-sided CDF ambiguity set with a fixed radius gives the exact worst-case expected latent rating maximum and an exact response-level policy credit for PPO. To preserve typical responses and limit policy drift, we impose lower bounds on the predicted mean rating and on an independently trained quality score, and limit changes from the reference policy using a standard KL measure of policy change with rollback. Across three seeds at , StableMax-PPO achieves a competitive robust maximum with lower held-out KL than Mean-PPO, group relative policy optimization (GRPO), and scalar max-PPO baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.