acceptodds
Under review as a conference paper at ICLR 2027

Reward Differences Matter for Robust LLM Alignment under Noisy Preferences

Abstract

Human preference alignment is pivotal for encouraging large language models to generate more helpful and safer responses that align with human values. Direct preference optimization (DPO) is a widely used alignment approach that enables LLMs to learn preferences directly from human preference data. However, human preference data inevitably contain noise arising from annotator inattention and bias, which can substantially degrade DPO performance. To identify noisy preferences, we conduct extensive preliminary experiments and empirically observe that as the reward difference decreases, the noise probability gradually increases during DPO training. Based on this observation, we propose a new method named (DACSDPO), which combines noisy preference correction with stability regularization. Specifically, DACSDPO first uses the reward difference to select and correct potentially noisy preferences during training. To further stabilize optimization in the presence of residual noise, DACSDPO then incorporates a stability penalty based on the squared norm of the sample gradient into the DPO objective, with an adaptive stability-aware coefficient. Experiments on Anthropic-HH, UltraFeedback, and PKU-SafeRLHF under varying noise rates demonstrate that DACSDPO improves robustness over representative preference optimization and robust alignment baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.