acceptodds
Under review as a conference paper at ICLR 2027

KKT-DPO: Direct Preference Optimization with Optimality-Derived Per-Prompt

Abstract

Direct Preference Optimization (DPO) aligns large language models (LLMs) with human preferences using a coefficient that controls deviation from a reference model but is fixed across prompts. Existing dynamic methods adjust across batches, pairs, or prompts using method-specific signals, rather than deriving it from the optimization problem. In this paper, we study a per-prompt KL-constrained formulation, in which each prompt has its own multiplier , and characterize the optimal through the KKT conditions. We show that the optimal multiplier decomposes into the product of the reward standard deviation under the reference model and a dimensionless factor determined by the standardized reward distribution and the KL budget. When prompts share the same standardized reward distribution, their optimal multipliers are proportional to their reward standard deviations. This suggests that the reward standard deviations determine how the optimal multipliers differ across prompts, and that an anchor sets the overall level. Based on this, we propose KKT-DPO, which approximates the relative reward standard deviations using responses sampled from the reference policy and scored by policy-to-reference log-ratios under a lagged policy snapshot. It then uses a coefficient tuned for vanilla DPO as a practical anchor to set the absolute multiplier. Experiments show that KKT-DPO attains the best performance among the baselines on the AlpacaEval and Arena-Hard and improves over DPO on reasoning, mathematics, coding, and instruction-following tasks, demonstrating that our per-prompt coefficients derived from the optimality conditions are effective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.