acceptodds
Under review as a conference paper at ICLR 2027

Query-Efficient Preference-based Robust Reinforcement Learning from Pairwise Comparisons

Abstract

We propose a preference-based robust reinforcement learning (PbRRL) framework that directly optimizes policies against the worst-case reward consistent with observed pairwise comparisons. The study focuses on settings in which only a very limited number of preference data are available and these data are highly trustworthy. Under the assumption of a linear reward model, the preference data with no response error induce a polyhedral ambiguity set over reward parameters, leading to a tractable max–min formulation of PbRRL. We propose a geometry-aware Vertical Cutting Method for a Cuboid, which adaptively generates preference queries to efficiently reduce the ambiguity set. We establish exponential convergence of the ambiguity set and asymptotic optimality of the resulting robust policy. Experiments on Ant-v5 show that our method substantially outperforms a traditional baseline when only a very limited number of preference queries are available. Our paper is among the first to bring the preference elicitation method that reduces the ambiguity set by pairwise comparisons directly into reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.