Beyond Unbiasedness: Conservative Reward Modeling from Observational Feedback
Abstract
Reward models (RMs) are central to reinforcement learning from human feedback (RLHF) for aligning large language models, while observational user feedback provides a scalable and deployment-aligned source of supervision. However, such feedback is selectively observed, inducing selection bias between the observed data and the target population. Existing propensity-based methods can correct this bias, but we show that unbiasedness alone does not guarantee reliable reward modeling. In particular, long-tailed observational feedback leaves low-support regions with near-zero propensities, where inverse weighting can amplify sparsely supported observations into spurious reward peaks that attract downstream policy optimization. To address this problem, we propose Conservative Reward Modeling (ConRM), which combines clipped propensity weighting with support-aware conservative regularization. Clipping limits the influence of near-zero-propensity samples, while the conservative regularizer uses well-supported anchors and local reward variations to construct data-supported intervals that constrain low-support extrapolation. Experiments on three public preference datasets under simulated long-tailed observational feedback demonstrate that ConRM consistently improves reward estimation across datasets and model backbones, while also yielding more reliable downstream policy alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.