acceptodds
Under review as a conference paper at ICLR 2027

GraphRubric: Learning Rubric Selection and Ordering for LLM-as-a-Judge

Abstract

Evaluation rubrics specify what an LLM judge should assess, yet prior work shows that rubric selection and ordering also affect its scores, even when prompts are semantically equivalent. We introduce GraphRubric, a lightweight graph controller conditioned on the input case that optimizes rubric presentation to improve agreement between judge scores and human labels. Rubrics are encoded as graph nodes that exchange information through graph attention; GRPO-style sampling converts within-group rewards into advantages for directed edges in their execution order. On SummEval, HANNA, and USR, three DeepSeek training seeds with fixed initialization reduce the primary error metric by 13.6%, 20.5%, and 5.8%, respectively, relative to uniform sampling over all nonempty paths. We also evaluate GLM and Gemini as judges. Across these datasets, GraphRubric shows generally favorable performance relative to the random and empty-rubric controls and the reported validation-selected fixed-path baselines. The method offers a prompt-engineering approach with potential applications to agent skill and harness optimization, reward-system construction, and user-data screening and cleaning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.