Auditing Length Confounds in DPO Data Selection
Abstract
Preference-data selection methods for Direct Preference Optimization (DPO) are often credited with improving alignment when they outperform standard or hard-negative baselines. Such comparisons can be confounded by response length: a selector may change the length relation between chosen and rejected responses rather than isolate better preference evidence. We present a shared-surface audit for length confounds in DPO data-selection studies. It requires a length-controlled baseline, common evaluation pairs, length-matched and slate-level metrics, threshold sweeps, and clustered uncertainty. Across controlled DPO runs over UltraFeedback TruthfulQA slates, soft length weighting, graph negatives, and exposure-aware selection initially improve over a hard-negative baseline, but their apparent wins become unstable under length control: one run reverses across all primary metrics, while another produces metric-dependent outcomes rather than a stable dominance relation. On a frozen RewardBench evaluation, standard DPO, length control, and exposure-aware selection obtain 50.8%, 58.3%, and 61.7% accuracy, respectively, but the latter's +3.3 point delta over length control has a 95% paired-bootstrap interval spanning zero. Additional held-out and generated-task evaluations show both supporting and countervailing cases. The contribution is not a universal selector; it is a falsification standard for deciding when DPO data-selection gains survive a shared, confound-aware audit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.