acceptodds
Under review as a conference paper at ICLR 2027

Plateau-DPO: Gradient-Envelope Allocation for Robust Preference Optimization

Abstract

Robust preference optimization must limit the influence of incorrect labels while preserving learning from valid preferences that the model has yet to learn. Both types of comparisons can have negative reference-relative margins, so downweighting comparisons based on their margins can also weaken useful supervision. For robust losses whose gradient weights concentrate around a single peak, valid comparisons away from that peak receive smaller weights even when they remain useful for learning. To address this issue, we introduce Plateau-DPO, which assigns a constant gradient weight to all comparisons within a prescribed margin interval and smoothly attenuates weights outside it. Comparisons within the interval are therefore not downweighted merely because their margins deviate from a peak, while extreme margins still receive reduced weights. We join the plateau with Gaussian tails and integrate the resulting weighting function to obtain a bounded symmetric preference loss. For a fixed interval, loss range, and integrated tail mass, we prove that the constant plateau maximizes the minimum gradient weight within the interval. Further analysis provides sufficient conditions for one-step descent of a target objective. Experiments under controlled label flips on Anthropic HH and UltraFeedback show improvements in preference accuracy and generation win rate. Gradient diagnostics and shape ablations show that the plateau retains larger normalized gradient weights on hard, unflipped comparisons, supporting the design of preserving learning signals across a finite margin interval. Code is available at https://anonymous.4open.science/r/Plateau-DPO-5B2F.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.