acceptodds
Under review as a conference paper at ICLR 2027

Towards Safety Alignment by Mitigating Imbalanced Preference Comprehension

Abstract

With the rapid development of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. Despite the effectiveness of DPO for safety alignment, prior studies have identified overoptimization as a persistent issue in this method, which can compromise safety generalization. This paper explores this phenomenon from the perspective of the model’s comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pair, which compromises the model’s safety generalization. To address this issue, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on an information-theoretic comprehension measure. Experiments across multiple LLMs demonstrate that B-DPO improves safety generalization over competitive preference optimization baselines while maintaining competitive general capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.