Flexible Quantile Reward Modeling for RLHF with Graded-Score Supervision
Abstract
Reward models are central to RLHF because they determine which model behaviors are reinforced during policy optimization. However, preference datasets with graded response scores are often reduced to binary chosen/rejected labels, discarding information about preference strength and reward location, while existing distributional approaches may impose restrictive parametric assumptions. We introduce the Flexible Quantile Reward Model (FQRM), which represents each response with a scalar-compatible center and an ordered quantile support capable of modeling skewed, asymmetric, and heavy-tailed reward distributions. Given comparable response-level scores, FQRM derives complementary supervision for relative ranking, preference strength, support location, and coarse reward regions, allowing the available graded information to constrain a single response-level representation. Unlike attribute-based quantile reward models, FQRM requires neither a predefined reward-attribute ontology nor attribute-level annotations. Experiments across multiple reward model benchmarks demonstrate the effectiveness of the proposed framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.