RubricRL: Rubric-Aware Reinforcement Learning for LLM Evaluation
Abstract
Reliable LLM judging requires not only learning , but also determining for each instance. Existing judge-training methods mainly optimize the former while leaving instance-specific evaluation criteria implicit, whereas rubric-guided approaches make these criteria explicit but typically optimize rubric generation separately from downstream judging. We introduce **RubricRL**, a reinforcement learning framework that couples instance-specific rubric learning with rubric-conditioned judging. RubricRL uses downstream judging performance to optimize instance-specific evaluation rubrics and, in turn, leverages the learned rubrics to guide score prediction within a shared policy model. We further introduce a distributional second-moment reward for absolute scoring quality and a distributional pairwise score-gap reward for relative scoring quality. Extensive experiments across diverse backbones and benchmarks show that RubricRL consistently improves LLM judging, achieves competitive performance with strong closed-source LLMs, and generates rubrics that provide judging performance comparable to human-validated criteria.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.