acceptodds
Under review as a conference paper at ICLR 2027

Not All Reward Matters: Uncertainty-Aware Reward Modeling for RLHF

Abstract

Reward models in reinforcement learning from human feedback (RLHF) distill human preferences into scalar signals that steer policy optimization, yet they offer no measure of evaluation confidence. Every input receives a single score, whether it is well-supported or ambiguous. Without uncertainty estimates, policy optimization methods such as Group Relative Policy Optimization (GRPO) treat all rewards as equally informative, propagating uncertain evaluations with the same weight as confident ones. To address this limitation, we propose **Uncertainty-Aware Reward Modeling (UARM)**, a framework that equips reward models with calibrated, per-sample uncertainty through quantile-based modeling and calibration. Each reward is represented as a point estimate and an adaptive prediction interval, enabling GRPO to downweight uncertain evaluations during advantage computation. Experiments demonstrate that UARM achieves higher reward prediction accuracy on uncertainty-ranked samples than baselines across three preference datasets, and delivers stronger downstream RLHF alignment. Code is available at https://anonymous.4open.science/r/UARM-42C0.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.