Do LLM Judges Faithfully Follow the Rubric? Measuring and Improving Rubric-Following in Pairwise Evaluation
Abstract
LLM judges are widely used to evaluate open-ended generation, and rubrics are increasingly employed to guide their evaluations using explicit criteria. This helps only if the judge faithfully follows the rubric, an assumption that is rarely tested. We test it on rubric-conditioned pairwise judges, which choose the better of two responses according to a given rubric. We find that when the rubric contradicts the judge's own preference, the judge ignores the rubric and keeps its verdict on up to 23.6% of judgments across four judges and four preference datasets. In most cases, the more confident the judge is before seeing the rubric, the more likely it is to ignore it. To address this, we propose RubricLens, which improves rubric-following by optimizing only the judge's instruction. For each pair of responses, it synthesizes two rubrics, each favoring a different response, so that one of them always conflicts with the judge's own preference. It then optimizes the instruction with GEPA to make the judge faithfully follow both rubrics. Across four judges and four preference datasets, RubricLens raises agreement with the rubric when it conflicts with the judge's own preference in every judge-dataset pair, from 86.2% to 93.7% on average, and lowers the rate at which the judge keeps its own choice to at most 12.0%. The same instruction improves rubric-following on expert-written rubrics without further tuning. RubricLens also outperforms fine-tuned rubric-conditioned judges in every setting, without changing model weights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.