acceptodds
Under review as a conference paper at ICLR 2027

RITR: Reliability-Informed Triage and Constrained Rewriting for Rubric-Based Reinforcement Learning

Abstract

Rubric-based reinforcement learning (RL) extends training to open-ended tasks, but standard importance-weighted aggregation ignores the reliability of grading evidence. Group normalization can turn judge fluctuations into unit-scale advantages when responses lack genuine quality differences, while uniformly failed criteria offer no successful examples. We propose RITR, a diagnose-before-repair framework that decomposes judge scores into decision direction and confidence. A criterion-level reliability proxy and four-state triage admit only sufficiently confident cross-boundary criteria to the main reward and control their relative weights. For criteria failed by all trusted responses, RITR allocates one sample from a fixed response budget to a targeted rewrite of the best-scoring response. The rewrite is gated on target satisfaction and preservation of previously satisfied criteria, and its scaffold is removed in two stages. Starting from Qwen3-8B, we train separate policies on medical, science, and finance rubric data. RITR improves over uniform-aggregation rubric RL on six of seven benchmarks, by up to 4.95 points on open-ended medical evaluation and 2.60 points on medical multiple-choice tasks. Medical ablations show open-ended gains from triage and further gains from constrained repair.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.