RAR²: Reliability-Aware Rubrics as Rewards for Reinforcement Learning
Abstract
LLMs are increasingly deployed across open-ended tasks, where high-quality outputs must satisfy multiple, task-specific criteria. Rubric-based evaluation makes these criteria explicit and is used to assess complex outputs in specialized, high-stakes domains. The Rubrics as Rewards (RaR) framework extends this approach to reinforcement learning by aggregating rubric-level verdicts into structured reward signals. We characterize rubric signals along two dimensions: discrimination across policy-generated responses and stability across repeated evaluations. The two dimensions divide rubric signals into four quadrants, with stable and discriminative signals providing the most reliable guidance for policy learning. However, standard RaR implicitly assumes noise-free rubric verdicts and therefore cannot distinguish genuine differences between responses from those induced by unstable judgments. To address this limitation, we propose Reliability-Aware Rubrics as Rewards (RAR²). RAR² measures stability through within-response variance and discrimination through between-response variance. A Beta–Binomial random-effects model then assesses whether each rubric's apparent discrimination is supported by stable judgments, and a reliability gate filters out rubric signals that lack such support. Experiments show that RAR² achieves the best performance across models of different sizes and diverse tasks. Additional analyses demonstrate that RAR² delivers consistent performance gains across judge reasoning-effort settings and improves agreement with expert assessments across judge models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.