acceptodds
Under review as a conference paper at ICLR 2027

When Does DPO Add Value over Continued SFT? Evidence from Correction-Preservation Frontiers

Abstract

After supervised fine-tuning (SFT), when does Direct Preference Optimization (DPO) add value over chosen-only continued SFT (CSFT)? Starting from shared SFT checkpoints, we track correction of inherited errors and preservation of initially correct decisions. Across two models and two ranking datasets, controlled increases in training disagreement widen the DPO-CSFT accuracy gap by 2.62–4.08 percentage points, mainly through CSFT decline rather than increased DPO gain over SFT. DPO incurs less corruption over overlapping correction ranges; the tested CSFT schedule and KL variants on MS MARCO do not recover its high-correction, low-corruption region. In naturally sampled ESCI and TL;DR, DPO instead gains mainly through additional correction despite greater corruption. These contrasting results show how DPO's relative value shifts between correction and preservation across training conditions. Exploratory analysis suggests that high training disagreement combined with a small SFT-error pool can signal stringent preservation requirements before further training; predicting realized gains remains open.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.