Sinkhorn-DPO: Distributionally Robust Preference Alignment via Entropic Optimal Transport
Abstract
Reinforcement learning from human feedback (RLHF) is a standard approach for aligning large language models (LLMs) with human preferences, while direct preference optimization (DPO) provides a simpler offline alternative that avoids explicit reward modeling. However, DPO remains sensitive to distribution shifts in preference data, motivating robust formulations based on distributionally robust optimization (DRO). Wasserstein-DPO (WDPO) methods provided a principled treatment of preference distribution shift, but their exact robust objectives involved nonsmooth dual structures that complicate gradient-based analysis and finite-sample estimation. We propose Sinkhorn-DPO (SDPO), a distributionally robust preference optimization framework based on entropically regularized optimal transport. Its dual formulation yields a smooth log-sum-exp envelope, enabling explicit derivative characterizations of the robust objective. Under mild regularity and a data-coverage condition imposed on the fixed reference measure, we prove smoothness and strong convexity of the profiled objective. Our finite-sample analysis establishes a parameter estimation rate of near- with sample size , which is a faster sample-size dependence than the bound of WDPO. The robustification is implemented through Sinkhorn-induced sample weights, so the policy is updated using a weighted first-order DPO gradient. We evaluate SDPO against DPO, Kullback-Leibler DPO (KLDPO), PO, and WDPO in linear experiments and in computational profiling with LLMs. The linear experiments examine radius calibration, parameter stability, and behavior under an induced within-corpus shift, while the LLM study reports runtime, peak GPU-memory usage and validation accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.