acceptodds
Under review as a conference paper at ICLR 2027

OrDiPO: Orthogonal Distillation for Preference Optimization in Language Models

Abstract

Knowledge distillation provides token-level teacher guidance that can complement pairwise preference learning for language model alignment. However, closer imitation of the teacher’s predictive distributions does not necessarily improve alignment with the target preferences. Through controlled gradient-component interventions, we show that the distillation residual orthogonal to an estimated preference-gradient direction retains useful generative guidance. Motivated by this observation, we propose Orthogonal Distillation for Preference Optimization (OrDiPO), which treats Direct Preference Optimization as the primary objective and chosen-prefix distillation as an auxiliary signal. Specifically, OrDiPO uses an exponential moving average of past preference gradients to define the projection normal. It removes the distillation-gradient component parallel to this normal and combines the remaining orthogonal residual with the unmodified current preference gradient. Experiments on UltraMedical-Preference and SHP–AskDocs demonstrate improvements over DPO and DPO+KD in both reference-corrected pair accuracy and generation win rate on internal test sets, with further evidence from external benchmarks. Additional experiments on HelpSteer3-Preference support applicability in a general-domain setting, while ablations and gradient analyses support the roles of orthogonal guidance and preference-gradient averaging.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.