Query-Efficient Preference-based Robust Reinforcement Learning from Pairwise Comparisons
Abstract
We propose a preference-based robust reinforcement learning (PbRRL) framework that directly optimizes policies against the worst-case reward consistent with observed pairwise comparisons. The study focuses on settings in which only a very limited number of preference data are available and these data are highly trustworthy. Under the assumption of a linear reward model, the preference data with no response error induce a polyhedral ambiguity set over reward parameters, leading to a tractable max–min formulation of PbRRL. We propose a geometry-aware Vertical Cutting Method for a Cuboid, which adaptively generates preference queries to efficiently reduce the ambiguity set. We establish exponential convergence of the ambiguity set and asymptotic optimality of the resulting robust policy. Experiments on Ant-v5 show that our method substantially outperforms a traditional baseline when only a very limited number of preference queries are available. Our paper is among the first to bring the preference elicitation method that reduces the ambiguity set by pairwise comparisons directly into reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.