acceptodds
Under review as a conference paper at ICLR 2027

Robust Nash Learning from Human Feedback against Preference Uncertainty

Abstract

Nash learning from human feedback (NLHF) models preference alignment as a two-player zero-sum game under general pairwise preferences directly, bypassing the restrictive latent reward assumption. But prior works typically assume this preference is unique and fixed. Yet in practice, such preference can be uncertain or pluralistic. We thus study robust NLHF, seeking a policy that is robust to all plausible preferences within some ambiguity set through optimizing the worst-case performance under the opponent policy and adversarial preference. However, the resulting double minimization does not generally preserve the convex-concave structure of standard NLHF. In this paper, we show that this difficulty can be removed exactly by lifting the adversary to a joint scenario-response law with conditional KL regularization, from which we restore the convex-concave structure and preserve the original robust value. We first consider the case of finite ambiguity sets, showing our reformulation yields a closed-form joint entropic Mirror-Prox algorithm with an convergence guarantee. We then extend the framework beyond finite scenario families to finite-dimensional compact-convex uncertainty sets with affine preference kernels. Although the natural continuous initialization has infinite KL radius against atomic worst cases, we use a localization argument to recover an rate. We further evaluate the performance and robustness of our framework. Our studies thus provide a unified optimization framework for robust alignment under uncertain preference beyond latent rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.