EvoJudge: Dual Evolution of Reusable Judging Skills and Rubrics for LLM-as-a-Judge
Abstract
LLM-as-a-judge plays an increasingly important role in model development, supporting fine-grained evaluation and providing reward signals for training. However, obtaining reliable judgments often requires highly capable frontier models, whose costs limit their use in large-scale evaluation and supervision. Inspired by recent advances in self-evolving agents, % which extract recurring patterns and convert them into reusable experience, in this paper, we introduce EvoJudge, a dual evolution framework that improves LLM judges along two complementary dimensions: how to judge, represented by procedure skills that guide evaluation strategies, and what to judge, represented by criterion rubrics that specify evaluation standards for each type. Starting from a few labeled samples, EvoJudge reflects on original judgment and synthesizes the reusable experience into skill and rubric libraries. In addition, periodic validation detects performance regressions, triggers selective rollback and provides corrective guidance for subsequent evolution. Experiments across model families demonstrate obvious judging improvements, with Qwen3-30B-A3B gains , , and on VerifyBench-hard, RewardBench2, and RMBench-hard, respectively, and enables Qwen3.5-35B-A3B to match frontier LLMs like GPT-5.6 Sol and Kimi-k3. Our analyses further demonstrate the complementary benefits of skills and rubrics, the effectiveness of periodic validation, and the transferability across various models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.