acceptodds
Under review as a conference paper at ICLR 2027

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

Abstract

Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objective does not specify how probability mass should be redistributed to achieve that margin. Consequently, preference separation can improve even as probability mass collapses away from both preferred and unrelated responses, a phenomenon associated with the squeezing effect. We show that this pathology is localized: continued negative updates become destructive when rejected responses enter low-probability valleys, where further suppression induces harmful redistribution through the coupled softmax geometry. We introduce Gradient-Gated Preference Optimization (Gate-DPO), which uses the current probability geometry of each rejected response to selectively attenuate its contribution, limiting excessive rejected suppression while delaying preference-loss saturation. Gate-DPO is objective-agnostic and composes with existing preference methods. Across four architectures, two preference datasets, and multiple objectives, gating consistently reduces squeezing and improves chosen-response likelihood, while exhibiting robust reductions in rejected-response collapse across architectures and objectives. Mass-dynamics analysis further shows that large preference margins can conceal substantial suppression of unrelated responses, while gating achieves healthier redistribution. Together, our results show that preference margin alone is insufficient to characterize optimization health and that selectively controlling probability dynamics provides a simple, general mechanism for stabilizing preference optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.