Learning Reward Distributions from Implicit Feedback with Sparse Expert Supervision
Abstract
Reward modeling is a key component of reinforcement learning from human feedback. Existing methods mainly rely on explicitly collected expert preferences, which provide high-quality supervision but are costly to obtain. In contrast, implicit feedback is abundant in real-world systems, but it is noisy and thus cannot be directly treated as expert supervision. Moreover, conventional reward models typically represent human preferences with a single expected reward, limiting their ability to capture complex variation in expert assessments. To address these challenges, we propose Reward Conditional Mean Embedding (RCME), a kernel-based framework that learns conditional expert-reward distributions from sparse expert preferences and abundant implicit feedback. Equipped with a characteristic kernel, RCME uniquely characterizes the conditional reward distribution and can be reused for different downstream reward utilities without retraining. We establish an identification result for RCME and develop a doubly robust estimator, which further show that the learning loss is Neyman-orthogonal to nuisance-estimation errors. For downstream policy optimization, we propose Group Kernel Relative Policy Optimization (GKRPO), which uses a kernel alignment score relative to a maximum-reward anchor and recovers mean-reward GRPO as a special case. We demonstrate the utility of our framework on synthetic datasets, PKU-Safety, and a real-world e-commerce application.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.