acceptodds
Under review as a conference paper at ICLR 2027

RubricLearner: Learning Evaluation Rubrics through Diagnosis-Guided Refinement

Abstract

As language models take on open-ended work such as deep research and office assistance, response quality hinges on many purpose-dependent requirements rather than a single correct answer, so evaluation depends on rubrics that state what a response must satisfy and how shortcomings affect its assessment. A rubric is nevertheless an empirical artifact: drafted criteria can miss the distinctions a task actually requires, and piecemeal score-driven edits risk discarding still-valid guidance while overfitting to the errors they fix. We introduce RubricLearner, a framework that improves evaluation rubrics through a diagnosis-guided refinement loop. Each iteration contrasts erroneous scoring justifications with reference judgments and distills conditional rules specifying when a lesson applies, what to attend to, and which mistake to avoid. Accepted rules accumulate, and the complete rubric is regenerated from them; a regenerated rubric is retained only when it improves validation performance. Because each rule states when it applies and accepted rules are retained rather than overwritten, the accumulated guidance stays specific to its context and grows across iterations. With Qwen3.7-Max as the judge, RubricLearner outperforms a state-of-the-art baseline across eight public benchmarks, an average relative gain of 3.51%. Beyond these scores, RubricLearner produces higher-quality rubrics with fewer defective criteria, and the improvement extends to unseen examples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.