Rubrics Learn from Near-Miss: Contrastive Anchoring for Rubric Evolution
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has driven strong post-training gains on tasks with verifiable answers, such as maths and code generation. Yet, extending the method to open-ended generation remains challenging, as ground-truth labels are often unavailable and human annotation is costly for these tasks. Rubric-based methods offer a promising alternative by decomposing response quality into explicit criteria. Recent approaches adapt these criteria during training, but adapted criteria often describe properties that most sampled responses already share, rather than what separates stronger responses from weaker ones. A central challenge is therefore to generate rewards that discriminate among the current policy's responses near this decision boundary, without labelled feedback. We propose Contrastive Anchoring for Rubric Evolution (), a self-improvement method for non-verifiable tasks that requires no human annotations, reference answers, or a stronger teacher model. derives criteria by contrasting a synthetic candidate response against a near-miss variant, surfacing properties that separate stronger from weaker responses on-policy at each point in training. outputs are preferred 68–73% of the time compared to the base model, with consistent gains across HealthBench, ResearchQA, and three models (Gemma-3-4B-it, Qwen3-4B, Qwen3-8B). On both benchmarks, CARE improves over the base model by 19–22%, more than doubling the gains of the strongest rubric baseline. We further show that the gains from on ResearchQA remain after excluding citation coverage, indicating improvements in other aspects of response quality. Overall, CARE shows how near-miss contrastive anchors can provide discriminative feedback to evolve rubrics on-policy for self-improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.