acceptodds
Under review as a conference paper at ICLR 2027

Understanding Reward Over-optimization in Rubric-based Reinforcement Learning

Abstract

Reinforcement Learning (RL) with rubric-based rewards has become a canonical framework for post-training large language models (LLMs) in open-ended domains. However, recent studies have shown that rubric-based RL remains susceptible to reward over-optimization, where the RL-trained policy exploits spurious patterns that satisfy the rubric without necessarily improving the underlying quality of its responses. We first demonstrate the prevalence of this phenomenon in medical and scientific tasks, observing substantial reward over-optimization under both rubric-based and rubric-free evaluation. Motivated by this observation, we systematically investigate the underlying causes through a series of controlled experiments. Our results rule out several potential explanations, showing that neither suboptimal reference guidance or limited coverage of diverse reference responses, nor the use of stronger judges for rubric construction, can fully account for the observed over-optimization. In contrast, we find that factors such as the number of rubrics associated with each prompt, or using preference-based reward with rubric guidance reward, can substantially mitigate reward over-optimization. Taken together, our findings highlight the importance of rubric quantity and reward design in understanding and mitigating reward over-optimization in rubric-based RL

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.