GSafePO: Geometric Optimizers for Reward-Safe Preference Learning
Abstract
Direct Preference Optimization (DPO) and related preference-learning methods optimize relative margins between chosen and rejected responses, but provide limited control over their absolute reward dynamics. This can lead to likelihood displacement, where chosen responses become less likely despite improved preference margins, and to reward over-optimization, where excessive absolute reward shifts drive unnecessary policy drift that can degrade generation quality even as the preference objective continues to improve. Recent methods address these effects largely through loss-level interventions. We introduce GSafePO (Geometrically Safe Preference Optimization), an optimizer framework that instead directly constrains the geometry of the realized parameter update. GSafePO projects the effective optimizer step onto a reward-safe set so that rewards for chosen responses do not decrease and rewards for rejected responses do not increase to first order, while deviating minimally from the unconstrained update. We characterize the low-dimensional geometry of such updates, derive conditions under which they remain aligned with descent on the original preference objective, and establish population-level guarantees for stochastic minibatch training. We instantiate the framework as two adaptive optimizers, GSafeRMSProp and GSafeAdam, extending the construction to adaptive preconditioned optimizers. Across multiple preference objectives and model configurations, GSafePO consistently converts undesirable common-mode reward motion into chosen-up/rejected-down dynamics while preserving competitive preference performance. Generation-quality evaluations further show that these corrected dynamics translate into significantly improved behavior in settings exhibiting likelihood displacement. These results establish update-geometry control as a modular mechanism for adding explicit reward safeguards to existing preference-learning objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.