Learning Robotic Skills from Offline Preference-based RL
Abstract
Preference-based reinforcement learning (PbRL) has become a promising framework for training agents in the absence of explicit reward functions, offering a principled approach to steer robot agents toward fulfilling human intentions. Existing methods typically follow a two-stage pipeline: learning a reward model from preferences, then applying off-the-shelf RL algorithms with the learned reward. However, since these learned proxy rewards are not the ground truth, their direct use in standard RL algorithms inevitably introduces compounding errors. To address this issue, we propose Preference-Score Actor-Critic (PSAC), an actor-critic framework that augments the critic update with a preference-score function derived directly from preference data, instead of relying solely on the learned reward model. Extensive experiments demonstrate that PSAC achieves state-of-the-art performance on 14 out of 17 D4RL benchmark tasks. Furthermore, we validate the efficacy of the proposed preference score augmentation through ablation studies, demonstrating significant performance improvements. Finally, using fewer than 700 manually labeled preference pairs, PSAC effectively controls a Unitree G1 robot to navigate rough terrain. Code and video demonstrations are available at https://anonymous.4open.science/r/PSAC-6686.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.