RGC: Global Reward Control for Rubric-Based LLM Post-Training
Abstract
Reinforcement learning (RL) post-training can improve factual accuracy, safety, tone, and other dimensions of model responses. The dimensions relevant to scoring a response vary across prompts, yet developers need global control over their contributions to training. We introduce Rubric-Grounded Coordinates (RGC), a framework for controlling behavior trade-offs through global reward weights. Developers define the dimensions and their global reward weights; prompt-specific rubrics score responses on the applicable dimensions. Our analysis shows how prompt applicability and variation among response scores alter the influence of these weights in group-relative policy optimization (GRPO). RGC uses a small set of training runs to calibrate how global reward weights affect scores across dimensions and guide Pareto search over these weights. We then train new models with the searched weights, achieving distinct behavior trade-offs and improvements on independent benchmarks. RGC achieves the best RubricBench accuracy among the evaluated judges, outperforming OpenRubric by 14.04 percentage points with the same backbone. On Qwen3-8B, our trained models achieve leading results on a broad set of representative public benchmarks among the compared training configurations; one of our trained models improves EQ-Bench3 by 144.8 Elo over training with prompt-specific rubric rewards. https://anonymous.4open.science/r/RGC-reward-system-86B2
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.