Preconditioned Direct Preference Optimization: Geometry and Acceleration
Abstract
Reinforcement learning from human feedback (RLHF) has become a central tool for aligning large language models with human preferences. Direct preference optimization (DPO) recasts RLHF as a policy optimization problem without explicitly estimating the reward function, and enhances the stability and efficiency of the training pipeline. Existing methods for DPO update based on policy gradients, and only the linear convergence rate is guaranteed. Therefore, to achieve acceleration, we consider using second-order information. A natural candidate is the Newton direction, which, however, involves the product of a Hessian inverse and the negative gradient, appearing computationally intractable in large-scale optimization. To circumvent this, we exploit the structure of DPO and, remarkably, characterize the regularized Newton direction via fully first-order information. This direction—which is used to propose a preconditioned method for DPO (PDPO)—encodes the geometry of the comparison graph, where each response is a node and each pairwise preference induces an edge. Consequently, we prove that updating along the Newton direction achieves quadratic convergence, and the updates along preconditioned directions restore the superlinear or the quadratic convergence rate under mild assumptions. Numerical experiments demonstrate that the proposed method achieves superior convergence rate and accuracy than the baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.