Nash Rejection Sampling: From Max–Min Games to On-Policy Optimization for Nash Learning from Human Feedback
Abstract
Reinforcement learning from human feedback (RLHF) is successful for large language models alignment, but the approaches commonly assume that human preferences admit a transitive Bradley–Terry representation. This assumption may fail for non-transitive preferences. Nash learning from human feedback (NLHF) addresses such general preferences by seeking a Nash equilibrium of the two-player constant-sum game induced by the preference relation. Existing approaches typically either require costly alternating updates, or adopt identity preference optimization (IPO)-style objectives. We introduce Nash Rejection Sampling (Nash-RS), an on-policy optimization framework that reduces NLHF to a standard RLHF-style policy-gradient problem. At each iteration, Nash-RS uses current-policy responses and preference feedback to rejection-sample opponents from a reference policy, yielding a distribution that provide implicit scalar rewards for the same on-policy rollouts. Consequently, Nash-RS can directly reuse existing on-policy RLHF algorithms, including PPO, while avoiding a max–min optimizer and its associated nested game updates. Compared with IPO-style approaches, Nash-RS performs on-policy updates. Theoretically, we establish linear last-iterate convergence of the exact update to the regularized Nash equilibrium, while the practical implementation approximates this update using finite rejection samples and neural policy optimization. Experiments across multiple preference datasets and language-model families compare Nash-RS with method-native NLHF baselines, demonstrating competitive performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.