acceptodds
Under review as a conference paper at ICLR 2027

DPO's Recovery Advantage Depends on Answer Presentation

Abstract

Direct preference optimization (DPO) can restore answers lost during supervised fine-tuning (SFT), but option order affects which failures are selected and how recovery is measured. We test whether DPO’s advantage over supervised continuation persists under a different option order. On MMLU-Pro, we select questions that Base answers correctly and SFT answers incorrectly, then evaluate unchanged SFT and both continuations on the same questions under both orders. Reordering alone lets unchanged SFT answer 29.6–34.1% of the selected failures correctly. On questions selected under the original order, DPO’s advantage falls from 8.72 to 2.58 percentage points for Qwen3-8B and from 9.53 to 1.46 for OLMo-2-7B. Gains in its relative advantage on the remaining questions conceal this decrease in the full-test comparison. Reversing the screening and evaluation orders gives the same pattern. Within Qwen3-8B’s recovery pools, DPO also changes more SFT answers, but a smaller fraction of those changes is correct. Recovery evaluation should therefore retest fixed questions across presentations and report how often each continuation changes an answer, alongside how often those changes succeed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.