acceptodds
Under review as a conference paper at ICLR 2027

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Abstract

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close responses create a bottleneck for policy optimization: when evaluator score gaps collapse or reverse, policy supervision becomes weak or misleading. We theoretically characterize why these gaps matter through a probability allocation view: shifting probability mass from one response to another yields a directional gain exactly equal to the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. We therefore propose DynamicRubric to co-evolve evaluators and policies using response-set-conditioned rubrics. An LLM can take either role and self-evolve by learning to supervise its own updates without an external reward model or judge. Using an 8B backbone, DynamicRubric produces an evaluator that is competitive with a 70B reward model on preference evaluation. On open-ended tasks, DynamicRubric provides stronger policy supervision than static rubrics generated by a 235B model, and performance gains generalize to verifiable tasks. A DynamicRubric-optimized model is fully deployed for AI answering in a large-scale production environment, handling tens of millions of requests per day and improving key online metrics. These results suggest a principle for post-training: evaluators should evolve with the policies they supervise.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.