RLVR: Reinforcement Learning with Verifiable Rubric-based Ranking
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with relatively well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements, where quality is typically specified by multi-dimensional rubrics rather than a single criterion. Since policy optimization generally consumes one scalar per rollout, rubric-based RL pipelines must map multiple criterion scores into a scalar reward. This aggregation step is often treated as score scaling, although it implicitly determines how different quality dimensions trade off during training. A prevailing practice normalizes each criterion and takes a linear aggregation, which implicitly assumes that cardinal score differences are comparable across criteria and that gains on one criterion can compensate for failures on another—assumptions that are unreliable when rubric criteria are semantically heterogeneous. We propose Reinforcement Learning with Verifiable Rubric-based Ranking (RLVR), a verifiable ranking paradigm for rubric-based RLVR. For each criterion, RLVR converts rubric scores into criterion-specific within-group ordinal outcomes and recovers a latent utility from the resulting within-group comparison matrix; these utilities are then converted into a single optimization signal for policy training. By retaining only within-group ordering and discarding raw score magnitudes, RLVR avoids calibrating heterogeneous rubric scales. RLVR further supports objective-preserving attribute adjustment: auxiliary attributes that are systematically associated with observed rankings but are not themselves training objectives can be incorporated into the estimation process without expanding the rubric or directly rewarding those attributes. Across three model scales and 16 benchmarks, RLVR consistently outperforms a broad range of representative rubric-based RLVR baselines, achieves the best overall performance on a majority of benchmarks at every scale. Further analysis shows that it controls systematic effects associated with reasoning efficiency and response formatting while preserving the underlying quality objective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.