More Optimization, Worse Alignment: Cycles and Safeguards in Iterative DPO
Abstract
Iterative preference learning fits a policy, adopts it as the next reference, and uses it to generate new comparisons. Better fitting appears desirable, but adoption changes the next optimization problem. We show that this feedback can turn deeper direct preference optimization (DPO) into a persistent reward loss. Under a fixed Bradley–Terry reward, we construct a regular nine-response policy family where one-step DPO converges globally, whereas 32-step DPO admits an attracting suboptimal two-cycle. Every frozen inner problem has globally convergent gradient descent, and the unique outer fixed point remains locally stable and reward-optimal within the family. Matched IPO converges globally at the same depth, scale, and learning rate, despite agreeing with DPO through second order at the optimum. We explain the separation by connecting the fourth loss derivative, comparison moments, and inner depth to the leading possible cubic difference. Even an arbitrarily small logistic component in an IPO mixture can produce cycles, with model-dependent depth and coexistence windows. A certified fifth-order analysis quantifies cycle onset and the GD step sizes needed to preserve coexistence. Finally, an IPO-envelope safeguard restores convergence in the product model by controlling the adopted displacement, while allowing arbitrary inner depths and exact minimization. Together, these results show why reliable iterative alignment requires control of policy adoption beyond accurate preference fitting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.