acceptodds
Under review as a conference paper at ICLR 2027

RubricRL: Rubric-Aware Reinforcement Learning for LLM Evaluation

Abstract

Reliable LLM judging requires not only learning , but also determining for each instance. Existing judge-training methods mainly optimize the former while leaving instance-specific evaluation criteria implicit, whereas rubric-guided approaches make these criteria explicit but typically optimize rubric generation separately from downstream judging. We introduce **RubricRL**, a reinforcement learning framework that couples instance-specific rubric learning with rubric-conditioned judging. RubricRL uses downstream judging performance to optimize instance-specific evaluation rubrics and, in turn, leverages the learned rubrics to guide score prediction within a shared policy model. We further introduce a distributional second-moment reward for absolute scoring quality and a distributional pairwise score-gap reward for relative scoring quality. Extensive experiments across diverse backbones and benchmarks show that RubricRL consistently improves LLM judging, achieves competitive performance with strong closed-source LLMs, and generates rubrics that provide judging performance comparable to human-validated criteria.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.