acceptodds
Under review as a conference paper at ICLR 2027

Beyond Uniform Forgetting in Alignment: Characterising, Predicting, and Recovering Retention in Sequential Preference Optimisation

Abstract

Sequential Direct Preference Optimisation (DPO) lets a model acquire new alignment objectives without retraining from scratch, but later training can degrade preferences learned earlier. It remains unclear whether such forgetting is uniform, predictable or recoverable, and how it should be measured. We study these questions across five preference settings (HH-RLHF, HelpSteer2, PKU-SafeRLHF, UltraFeedback and Summarize-Feedback), organised as , and . When measured relative to the original base model, earlier objectives show degradation in some settings and stability or improvement in others. Across 16 retention comparisons, comprising ten two-stage orderings and six three-stage steps, the policy's own chosen-over-rejected ranking is substantially more stable than the reference-relative metrics. For length-normalised accuracy, a decline of more than 2.5 percentage points is ruled out in 13 of 16 comparisons, so large reference-relative declines need not indicate preference reversals. Before Stage 2, the Stage-1 model's margin on the incoming objective matches the direction of later margin retention in nine of ten orderings. The evidence comes from five datasets and describes margin changes rather than the policy's own accuracy. Replaying a small share of earlier data shows nominal improvements in margin retention in seven of nine orderings, while policy accuracy improves in only two.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.