Reinforced DPO: Single-Stage Preference Optimization Provably Equivalent to RLHF
Abstract
Direct preference optimization (DPO) aligns language models in a single training stage, offering a simpler and more stable alternative to the two-stage pipeline of reinforcement learning from human feedback (RLHF). Yet despite being derived from the RLHF objective, DPO fails to reproduce the desirable properties of RLHF: it learns only from offline preference data without on-policy learning, and it does not effectively regularize the policy toward the reference model. In this paper, we first prove that DPO provably deviates from RLHF in both respects. We then trace this inequivalence to a gap in DPO's derivation: the change of variables from rewards to Q-functions silently drops the reward-boundedness constraint that makes RLHF's KL-regularized RL stage well posed. Reinstating this constraint, which takes the form of a Bellman constraint on Q-functions, yields Reinforced Direct Preference Optimization (ReDPO). We prove that ReDPO is equivalent to RLHF and thus inherits its bound on the divergence from the reference policy. Crucially, the Bellman constraint induces a value flow that propagates both the successor Q-values and the reference log-probability into the current Q-value, carrying out KL-regularized RL through dynamic programming. Experiments show that ReDPO clearly outperforms DPO and its variants, while keeping the policy closer to the reference model and mitigating the forgetting of general capabilities. To our knowledge, ReDPO is the first single-stage preference optimization method that is provably equivalent to RLHF.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.