Robust Nash Learning from Human Feedback against Preference Uncertainty
Abstract
Nash learning from human feedback (NLHF) models preference alignment as a two-player zero-sum game under general pairwise preferences directly, bypassing the restrictive latent reward assumption. But prior works typically assume this preference is unique and fixed. Yet in practice, such preference can be uncertain or pluralistic. We thus study robust NLHF, seeking a policy that is robust to all plausible preferences within some ambiguity set through optimizing the worst-case performance under the opponent policy and adversarial preference. However, the resulting double minimization does not generally preserve the convex-concave structure of standard NLHF. In this paper, we show that this difficulty can be removed exactly by lifting the adversary to a joint scenario-response law with conditional KL regularization, from which we restore the convex-concave structure and preserve the original robust value. We first consider the case of finite ambiguity sets, showing our reformulation yields a closed-form joint entropic Mirror-Prox algorithm with an convergence guarantee. We then extend the framework beyond finite scenario families to finite-dimensional compact-convex uncertainty sets with affine preference kernels. Although the natural continuous initialization has infinite KL radius against atomic worst cases, we use a localization argument to recover an rate. We further evaluate the performance and robustness of our framework. Our studies thus provide a unified optimization framework for robust alignment under uncertain preference beyond latent rewards.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.