acceptodds
Under review as a conference paper at ICLR 2027

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Abstract

Reinforcement Learning with Rubric Rewards (RLRR) is a framework that extends conventional reinforcement learning from human feedback (RLHF) and verifiable rewards (RLVR) by replacing scalar preference signals with structured, multi-dimensional contextual rubric-based evaluations. However, existing approaches in RLRR are limited to linearly compress vector rewards into a scalar reward with fixed weightings, which is sensitive to artificial score design and fails to capture correlations among reward dimensions. To overcome the limitations of reward aggregation, this work proposes Alternating Reinforcement Learning with Rubric Rewards (ARL-RR), a framework that removes the need for a fixed scalarization by optimizing one semantic rubric meta-class at a time. Theoretically, we show a variance contraction effect of reward aggregation that explains the performance improvement. We further introduce a lightweight search-based adaptation procedure that selects the next meta-class dynamically based on task performance, enabling the policy to emphasize critical objectives and thus improve the model performance. Empirically, our experiments on the HealthBench dataset with experts' annotation demonstrate that ARL-RR uniformly outperforms scalarized methods in both model performance and training efficiency across different model scales (1.7B, 3B, 8B, and 14B).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.