Preference Learning Beyond Confidence: Efficient Gradient-Based RLHF via Adversarial Preferences
Abstract
Recent advances in Large Language Models (LLMs) and robotics have intensified interest in Reinforcement Learning from Human Feedback (RLHF), where reward functions are learned from pairwise human preferences modeled via sigmoid-based utility functions. Despite their widespread adoption, existing RLHF frameworks either lack theoretical guarantees or are computationally intractable — relying on confidence-set-based methods that cannot scale to continuous, high-dimensional decision spaces. We address this gap by developing Mirror-Duel, an efficient gradient descent-based algorithm with provably optimal regret guarantees under adversarially changing preferences, without assuming a fixed stochastic reward model. Formalizing the problem as an adversarial online convex optimization (OCO) problem with only weak pairwise preference feedback, our online mirror descent (OMD) approach achieves an optimal regret bound — matching a lower bound — with only runtime. Beyond pairwise preferences, we extend Mirror-Duel to -batched and top- partial ranking feedback, respectively yielding improved and regret guarantees. We validate our theoretical findings extensively on synthetic adversarial instances as well as real-world RL environments (CartPole-v1, Acrobot-v1), demonstrating that Mirror-Duel consistently outperforms UCB-based baselines, including Online-IPO, APO, and Nash-MD in regret performance. Our work lays a theoretical and algorithmic foundation for scalable, gradient-based reward learning from human preferences, contributing to the broader effort of improving AI alignment with Human Preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.