RARE: Reinforcement Learning with Adaptive Rubric Evolution for Open-Ended Generation
Abstract
Verifiable rewards have advanced reinforcement learning in mathematics and code. However, many valuable open-ended tasks share three properties: no automatic verifier of overall quality, multiple requirements to satisfy jointly, and multiple valid responses to each query. Training therefore needs feedback on requirement satisfaction across different valid responses. Rubric-based reinforcement learning provides this feedback by aggregating scores from explicit criteria. As learning progresses unevenly, some criteria become fully satisfied and lose discrimination while others remain informative. Dynamic methods revise shared rubrics or select evolving criteria to restore discrimination, but may also change useful criteria. To address this, we propose Reinforcement Learning with Adaptive Rubric Evolution (RARE). The core idea is to use current policy feedback to guide both which criteria change and what their replacements assess. Sample-level updates replace criteria fully satisfied across current outputs with checks of remaining quality differences, retaining all other criteria and scores. Within this rule, we compare rewarding desired qualities with penalizing observed flaws. Video script generation exemplifies these properties and supports content production. Its demands for coherent, executable narratives under user constraints make it a challenging testbed. Targeting flaws performs best across three video types with a mean gain of 3.82 points over matched static rubrics. Model and human preferences support these gains. Improvements also extend to writing and general generation benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.