acceptodds
Under review as a conference paper at ICLR 2027

When Higher Scores Yield Less Feedback: Score Saturation in Self-Rewarding Language Models

Abstract

Self-rewarding language models can improve through an iterative feedback loop in which the model generates multiple responses, evaluates them using a discrete scoring rubric, and trains on the resulting preferences. A known failure mode is , where generated responses increasingly receive scores near the top of the rubric, producing more ties and fewer informative comparisons. We develop a theoretical analysis of when and how preference optimization contributes to saturation and reshapes the feedback available to subsequent rounds. We first show that the preference optimization update direction inherently contains a component that increases the probability of the highest score, while other components of the preference gradient can either reinforce or counteract this effect. We further decompose score changes into effects arising from changes in the model's response distribution and changes in its evaluation behavior, and derive conditions under which these effects accumulate to produce score saturation. As saturation progresses, the preference comparisons increasingly reflect the gradient structure associated with increasing top-score probability, even as informative comparisons become increasingly rare. Consequently, maintaining a fixed number of them requires a sampling budget that grows inversely with the remaining probability of a non-top score, and asymptotically exponentially in the accumulated gains. Experiments on self-rewarding language models validate these theoretical results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.