acceptodds
Under review as a conference paper at ICLR 2027

Flexible Quantile Reward Modeling for RLHF with Graded-Score Supervision

Abstract

Reward models are central to RLHF because they determine which model behaviors are reinforced during policy optimization. However, preference datasets with graded response scores are often reduced to binary chosen/rejected labels, discarding information about preference strength and reward location, while existing distributional approaches may impose restrictive parametric assumptions. We introduce the Flexible Quantile Reward Model (FQRM), which represents each response with a scalar-compatible center and an ordered quantile support capable of modeling skewed, asymmetric, and heavy-tailed reward distributions. Given comparable response-level scores, FQRM derives complementary supervision for relative ranking, preference strength, support location, and coarse reward regions, allowing the available graded information to constrain a single response-level representation. Unlike attribute-based quantile reward models, FQRM requires neither a predefined reward-attribute ontology nor attribute-level annotations. Experiments across multiple reward model benchmarks demonstrate the effectiveness of the proposed framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.