acceptodds
Under review as a conference paper at ICLR 2027

BRM: Better Reward Modeling via Critic Feature Transfer in Continuous Control

Abstract

Preference-based Reinforcement Learning (PbRL) replaces hand-crafted reward functions by learning a reward model from pairwise trajectory preferences, achieving good performance in several continuous control tasks. However, the generalization of the reward model quickly becomes a bottleneck of the pipeline: it has to predict rewards for the millions of visited state-action pairs from a few hundred preference trajectories. This data asymmetry results in the reward model overfitting to the labeled data, ultimately degrading final performance. In this work, we address this issue by turning this data asymmetry into an opportunity. We propose BRM (Better Reward Modeling), which copies the features of the critic, trained on the full replay buffer, and uses them as an input to the reward model, trained on only a few hundred labeled preferences. Across twelve tasks from the DeepMind Control Suite and MetaWorld and against seven PbRL baselines, BRM reaches a task-normalized Interquartile Mean (IQM) of 0.85 versus 0.55 for the strongest of them. Our empirical analysis indicates that BRM's gains are consistent with better reward generalization, most clearly on trajectories held out from the preference distribution. Our results suggest that the critic the agent already trains is an underused starting point for the reward model in existing PbRL methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.