acceptodds
Under review as a conference paper at ICLR 2027

Population-Relative RLHF for LLM Output Diversity: A Bi-Level Optimization Approach

Abstract

We study the problem of aligning large language models (LLMs) when output diversity is critical alongside output quality. Standard alignment methods (implicitly) optimize a reward model that scores individual responses in isolation, which can lead to mode collapse and overly concentrated policies. We propose population-relative RLHF, a framework that conditions pairwise preference feedback on a reference population of responses, enabling the reward model to capture a holistic quality-diversity trade-off. We formulate population-relative RLHF as a bi-level optimization problem, where the upper level learns a population-relative reward model and the lower level solves for the best-response LLM policy induced by that reward. We show that this formulation is an instance of a broader class of bi-level problems whose lower level constitutes a mean-field-type stochastic game. For this general problem class, we develop a first-order Hessian-free algorithm with provable finite-time convergence guarantees. We apply this algorithm to population-relative RLHF, and also introduce a simplified, memory-efficient variant suitable for large-scale implementation. Empirical evaluations demonstrate that our methods outperform prior art on the quality–diversity trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.