EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics
Abstract
Language models encode substantial evaluative knowledge from pretraining, yet current post-training methods rely on external supervision (human annotations, proprietary models, or scalar reward models) to produce reward signals. Each source brings a different constraint: human judgment is costly and hard to scale, proprietary APIs create dependencies, and verifiable rewards cover only domains with ground-truth answers. A model’s own evolution offers a natural source of supervision for self-improvement, yet remains largely untapped by current methods. We introduce EVOLM, a post-training method that turns this supervision into rewards for further learning. EVOLM trains two capabilities within a single language model in alternation: (1) a rubric generator producing instance-specific evaluation criteria optimized for discriminative utility, which maximizes a small frozen judge’s ability to distinguish preferred from dispreferred responses; and (2) a policy trained using those rubric-conditioned scores as reward. All preference signals are constructed from the policy’s own outputs via temporal contrast with earlier checkpoints, requiring no human annotation or external supervision. The co-trained policy achieves 69.2% average on the OLMo3-Adapt suite, exceeding policies trained with GPT-4.1-prompted rubrics by 2.5 points and those trained with the state-of-the-art 8B reward model Skywork-RM-V2 by 9.5 points. The learned rubrics transfer to judges they were never trained against, raising a Qwen3-8B judge’s RewardBench-2 accuracy from 39.7% to 62.4%. Rubrics that co-evolve with the policy therefore supply training signal that external reward sources do not.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.