When Does DPO Add Value over Continued SFT? Evidence from Correction-Preservation Frontiers
Abstract
After supervised fine-tuning (SFT), when does Direct Preference Optimization (DPO) add value over chosen-only continued SFT (CSFT)? Starting from shared SFT checkpoints, we track correction of inherited errors and preservation of initially correct decisions. Across two models and two ranking datasets, controlled increases in training disagreement widen the DPO-CSFT accuracy gap by 2.62–4.08 percentage points, mainly through CSFT decline rather than increased DPO gain over SFT. DPO incurs less corruption over overlapping correction ranges; the tested CSFT schedule and KL variants on MS MARCO do not recover its high-correction, low-corruption region. In naturally sampled ESCI and TL;DR, DPO instead gains mainly through additional correction despite greater corruption. These contrasting results show how DPO's relative value shifts between correction and preservation across training conditions. Exploratory analysis suggests that high training disagreement combined with a small SFT-error pool can signal stringent preservation requirements before further training; predicting realized gains remains open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.