acceptodds
Under review as a conference paper at ICLR 2027

Efficient Exploration for Iterative Nash Preference Optimization

Abstract

Preference alignment is central to improving large language models, but standard reward-based formulations can be restrictive when human preferences are cyclic, non-transitive. Nash Learning from Human Feedback (NLHF) addresses this limitation by modeling alignment as a preference game and targeting a Nash equilibrium. However, the learning-theoretic foundations of scalable NLHF remain limited. Existing regret guarantees requires training a general preference model and solving minimax oracles, while iterative NLHF methods are easier to implement but lack regret guarantees. In this paper, we study online iterative NLHF and identify exploration as the key obstacle. First, we show that standard iterative NLHF can suffer an exponential dependence on the KL-regularization parameter, revealing that implicit exploration through policy updates is insufficient for efficient exploration. Second, we propose Exploratory Nash Policy Optimization (ENPO) that combines SFT-based regularization with adversarial policy exploration. ENPO eliminates the exponential dependency without using any minimax oracles and explicit preference model estimations. Additionally, we present another theoretical algorithm: Bonus-explorer ENPO (BENPO) with an explicit exploration mechanism, which uses additionally oracles but can guarantee an regret bound. A practical algorithm: Direct ENPO (DENPO) was implemented for LLM fine-tuning. The empirical demonstrations show that DENPO consistently outperforms existing similar RLHF and NLHF baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.