Population-Relative RLHF for LLM Output Diversity: A Bi-Level Optimization Approach
Abstract
We study the problem of aligning large language models (LLMs) when output diversity is critical alongside output quality. Standard alignment methods (implicitly) optimize a reward model that scores individual responses in isolation, which can lead to mode collapse and overly concentrated policies. We propose population-relative RLHF, a framework that conditions pairwise preference feedback on a reference population of responses, enabling the reward model to capture a holistic quality-diversity trade-off. We formulate population-relative RLHF as a bi-level optimization problem, where the upper level learns a population-relative reward model and the lower level solves for the best-response LLM policy induced by that reward. We show that this formulation is an instance of a broader class of bi-level problems whose lower level constitutes a mean-field-type stochastic game. For this general problem class, we develop a first-order Hessian-free algorithm with provable finite-time convergence guarantees. We apply this algorithm to population-relative RLHF, and also introduce a simplified, memory-efficient variant suitable for large-scale implementation. Empirical evaluations demonstrate that our methods outperform prior art on the quality–diversity trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.