DPO's Recovery Advantage Depends on Answer Presentation
Abstract
Direct preference optimization (DPO) can restore answers lost during supervised fine-tuning (SFT), but option order affects which failures are selected and how recovery is measured. We test whether DPO’s advantage over supervised continuation persists under a different option order. On MMLU-Pro, we select questions that Base answers correctly and SFT answers incorrectly, then evaluate unchanged SFT and both continuations on the same questions under both orders. Reordering alone lets unchanged SFT answer 29.6–34.1% of the selected failures correctly. On questions selected under the original order, DPO’s advantage falls from 8.72 to 2.58 percentage points for Qwen3-8B and from 9.53 to 1.46 for OLMo-2-7B. Gains in its relative advantage on the remaining questions conceal this decrease in the full-test comparison. Reversing the screening and evaluation orders gives the same pattern. Within Qwen3-8B’s recovery pools, DPO also changes more SFT answers, but a smaller fraction of those changes is correct. Recovery evaluation should therefore retest fixed questions across presentations and report how often each continuation changes an answer, alongside how often those changes succeed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.