acceptodds
Under review as a conference paper at ICLR 2027

Perfect Preference Fit, Wrong Policy: Limits of DPO under Aggregation

Abstract

Can a policy fit every observed preference perfectly and still allocate its probability mass incorrectly? We study population direct preference optimization (DPO) after response identities are replaced by semantic classes, using the class allocation from fine-policy optimization as the benchmark. Sparse comparison designs separate three guarantees: fitting preferences, repairing one environment, and preserving the target across environments. On paths, small within-class variation produces sharp regret despite exact fitting and optimal choices of positive weights, temperature, and regular proper loss. This rate persists under arbitrary small offsets to the repeated class means. Closing the path restores preservation at the symmetric baseline, but nearby environments expose a gap: each admits exact recovery by balanced weights, while every shared balanced weighting has a worst-case lower bound even with a separately optimized temperature. Equal weights attain a matching upper bound. A variance-factored link representation extends the cycle rates to every sufficiently small-support common reward law, uniformly in graph size and atom masses. The benchmark measures loss of a KL-regularized objective; it can be positive with unchanged expected reward. Tree and dense-graph upper bounds give complementary positive cases, while sharp prediction on general sparse graphs remains open.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.