acceptodds
Under review as a conference paper at ICLR 2027

Conformal Rubrics: Calibrating Rubric Rewards with Human Score for Policy Learning

Abstract

Rubrics provide explicit criteria, such as correctness and helpfulness, for evaluating language model responses and constructing rewards for policy learning. Recent methods improve these rewards by refining the criteria or adjusting how their scores are combined. However, specifying what to evaluate does not ensure that a judge model assigns scores consistent with human judgments. Human ratings provide a reference for correcting these scores, but collecting ratings for every new response during policy training would be costly. We therefore propose Conformal Rubrics, which use a fixed human-rated dataset to calibrate rubric rewards for downstream policy training. The key challenge is estimating how humans would score responses on new tasks, where queries differ from those in the annotated dataset. We train a predictor on the source ratings and use its prediction errors to construct a conformal interval to include possible human scores for each rubric with a pre-specified coverage level. We then calibrate judge scores against these intervals by clipping scores outside them to the nearest boundary and retaining those inside. The corrected scores are then combined into rewards using existing rubric-based methods. Experiments show that calibration with the human-rated dataset HelpSteer reduces judges' rubric scoring errors and improves policy performance on different datasets, such as HelpSteer2, WildChat, and AlpacaEval-2. Our code is available at anonymous GitHub.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.