Reusing Preference Data Changes What DPO Learns
Abstract
Direct preference optimization (DPO) is often run for several passes over a smaller preference pool. We show that, at a fixed training budget, this reuse changes the direction of what the model learns. We compare fresh and reused data with identical steps, batch sizes, and pair presentations, measure each trained model’s change from its parent on a frozen 200-scenario battery and 10,273 held-out human-preference pairs, and calibrate every comparison against variation between training seeds. In OLMo-2 1B, reuse and fresh training agree at a reliability-corrected r* = 0.798 [0.675, 0.892] (about 37°), outside seed-to-seed variation, while a second, disjoint reused pool lands in the same place (r* = 0.966 with the first). At a fixed budget, the turn grows with the number of passes, from 20° at two passes to 38° at 6.4, as predicted by a finite-pool model of DPO, and a third seed confirms the effect (r* = 0.770 [0.637, 0.881]). The redirection replicates across four DPO model families and on held-out human preferences. It also reaches behavior: reuse changes the preferred response on one-fifth of held-out human-preference pairs, five to eight times the seed-to-seed rate, without improving agreement with human labels, and no comparison finds its generations better; at greedy decoding, fresh training’s response wins on 71–74% of prompts. Held-out preference accuracy is also lower under reuse (0.604/0.616 versus 0.651/0.660). How often each pair is reused is therefore part of a preference-training recipe: equal budgets can produce systematically different models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.