Rubrics Beyond Evaluation:Exploiting Rubric Structure in Reinforcement Learning
Abstract
Rubrics are widely used to evaluate open-ended question answering by breaking overall quality into explicit criteria (Kim et al., 2024; Arora et al., 2025). These explicit criteria make rubrics a rich source of learning guidance, yet their potential for guiding the training process has been largely overlooked. Existing methods use rubrics mainly as direct feedback signals, without explicitly modeling the capabilities behind their criteria (Gunjal et al., 2025; Li et al., 2026). This leaves their capability structure largely unused in guiding the training process. We introduce Rubric-Structured Reinforcement Learning (RSRL), a RL framework that unlocks the full potential of rubrics by exploiting their structure throughout the training process, rather than treating them simply as direct reward signals. At its core, RSRL builds on three key designes: (1) Rubric-to-Capability Structuring first recovers the capability structure hidden in rubric criteria, (2) SelfSufficient Capability Learning then uses this structure to bootstrap underdeveloped capabilities into effective learning, (3) Capability-Residual Credit Routing finally preserves capability-specific gains during optimization. Experiments show that RSRL consistently improves both overall performance and underdeveloped capabilities, boosting downstream scientific question-answering performance by up to 7.43 over the initial policy and 7.60 over non-rubric structured RL baselines, showing that reinforcement learning can substantially benefit from the capability structure encoded in rubrics
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.