Variance-Reduced Off-Policy Preference Optimization via Regularizer Correction
Abstract
Direct Preference Optimization (DPO) regularizes a learned policy toward a reference policy using an implicit reverse KL divergence. Other convex regularizers are also well motivated for preference optimization but are known to perform worse in practice – especially under off-policy training where the preference data does not come from the policy currently being optimized. Moving beyond the bandit setting to multi-step MDPs, this effect gets exacerbated by the fact that constructing an implicit reward function for convex regularizers in multi-step MDPs requires sampling subsequent actions from the current policy, causing estimation errors to compound along trajectories. We formally characterize the regularizers that do not require subsequent action sampling and prove that the reverse KL is unique in providing a closed-form solution. To address error compounding in convex regularizers other than the reverse KL, we propose a regularity condition that minimizes the variance of the implicit reward estimator and derive a canonical form satisfying this condition by design. Our experimental results on tabular MDPs and for large language model post-training show that the proposed correction consistently improves on alternative approaches to estimating separable convex regularizers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.