Towards Safety Alignment by Mitigating Imbalanced Preference Comprehension
Abstract
With the rapid development of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. Despite the effectiveness of DPO for safety alignment, prior studies have identified overoptimization as a persistent issue in this method, which can compromise safety generalization. This paper explores this phenomenon from the perspective of the model’s comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pair, which compromises the model’s safety generalization. To address this issue, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on an information-theoretic comprehension measure. Experiments across multiple LLMs demonstrate that B-DPO improves safety generalization over competitive preference optimization baselines while maintaining competitive general capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.