Active Preference Querying for Offline-to-Online Reinforcement Learning
Abstract
Offline-to-online reinforcement learning usually assumes scalar rewards during online fine-tuning. We study offline-to-online preference-based RL, where an agent starts from offline preference data and can improve online only through a limited budget of pairwise preference queries, so deciding which comparisons to label is central. We propose an active preference-query rule that selects comparisons that are ambiguous to a reward ensemble and informative for the Bradley-Terry reward estimator, and APOLLO, an algorithm that applies this rule to long-horizon control with offline initialization and a flow-based, action-chunked policy. Our analysis gives a finite-sample bound on the preference-prediction error that holds for adaptively chosen queries; the query rule enters only through how well the labeled data cover the policy's comparisons, so its guarantee improves on random querying when the rule improves this coverage. On eight OGBench manipulation tasks, APOLLO outperforms five preference-based and behavioral-cloning baselines. When only the query rule is changed, our rule achieves the best average and worst-case success among six rules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.