RubricEvo: Sustaining Learning Signals via Evolving Graph-based Multi-Level Rubrics in Open-Ended Reinforcement Learning
Abstract
Rubric-based reinforcement learning provides a practical way to optimize open-ended tasks without verifiable answers, yet fixed, flat rubrics become increasingly ineffective as the language-model policy evolves. Once broad criteria are mastered, rubric scores saturate and reward gaps among on-policy responses diminish; simultaneously, optimizing against a single global checklist favors a narrow set of rubric-compatible behaviors and suppresses alternative valid strategies. We propose , which represents each query-specific rubric as an evolving, multi-level graph. The graph organizes rubric into fixed nodes that preserve fundamental requirements, nodes that capture major quality dimensions, and nodes that provide fine-grained, strategy-dependent guidance. While Principle nodes remain unchanged, Aspect and Specific nodes can be replaced or branched as the policy evolves. During training, RubricEvo measures reward resolution and semantic diversity over on-policy response groups to guide the evolution of the graph-based rubric. Experiments across multiple domains and Qwen3-1.7B/4B/8B policies show that RubricEvo improves downstream performance while preserving more discriminative rewards and greater high-quality output diversity. All datasets and code are available at https://anonymous.4open.science/r/RubricEvo/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.