acceptodds
Under review as a conference paper at ICLR 2027

Reward Model Overoptimisation in Iterated RLHF

Abstract

Reinforcement learning from human feedback (RLHF) is a widely used method for aligning large language models with human preferences. However, RLHF often suffers from reward model overoptimisation, in which models overfit to the reward function, resulting in non-generalisable policies that exploit its idiosyncrasies. A common mitigation is *iterated RLHF*, in which reward models are repeatedly retrained with updated human feedback and policies are re-optimised. Despite its increasing adoption, the dynamics of overoptimisation in this setting remain poorly understood. In this work, we evaluate several iterated RLHF pipelines and measure their effect on overoptimisation. In particular, we systematically analyse how reward model training data is transferred across iterations, which reward function is used for optimisation, and how policies are initialised. Using the controlled AlpacaFarm benchmark with policies of M and B parameters and comparing the distributions of proxy and gold rewards, we observe that overoptimisation decreases over successive iterations as reward models increasingly approximate ground-truth preferences, although the gains shrink in later iterations. Reinitialising the policy from the supervised fine-tuned model at every iteration is the most robust choice. Continuing from the previous policy works as long as that policy has not overoptimised, which larger reward models make possible, while an overoptimised policy is hard to recover. Concatenating the preference data of all iterations is the most effective mitigation, and together with policy resets it is the pipeline we recommend.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.