acceptodds
Under review as a conference paper at ICLR 2027

The Real Conflict of Multi-objective Alignment Is Hidden in the Preference Labels

Abstract

Multi-objective alignment seeks to reconcile conflicts among helpfulness, safety, and honesty. However, our experiments show that these conflicts can depend on the training signal: they emerge in DPO training on preference data, whereas training guided by explicit reward functions can improve these qualities simultaneously. This discrepancy points to a hidden preference conflict: learning harmful compliance or sycophancy encouraged by the labels can undermine safety or honesty. These behavioral incentives are implicit in the preference data rather than specified as separate rewards. We propose Conflict-Disentangled Direct Preference Optimization (CD-DPO), which learns a weighted combination of frozen reward-model scores from examples annotated for these behaviors and uses it to correct preference labels before standard DPO training. Allowing negative weights lets the method identify unwanted behaviors through differences between reward scores rather than simply adding them. In a stylized model, we prove that allowing negative weights can suppress these behaviors more effectively than nonnegative reward scalarization at the same KL budget, and quantify the effects of fitting and training errors. Compared with vanilla DPO, CD-DPO achieves large safety gains on Qwen2.5 (0.5B, 3B, and 7B) and improves honesty and truthfulness on both Qwen2.5 and Mistral-7B at the same policy-training budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.